Staff Backend Engineer (Application Core Services)
Grafana Labs · United States
About The Role
Join Grafana, a leading open-source analytics and monitoring platform. As a Staff Backend Engineer, you will work on critical systems that power our business and directly impact our customers. You will design, build, and operate reconciliation systems, improve operational efficiency, and collaborate with various teams. Enjoy a remote-first work environment, competitive benefits, and opportunities for professional development. - Design, build, and operate reconciliation systems, including the SSS backend, to track desired stack state, detect and repair drift across stack templates, state, Hosted Grafana, and actual customer stack configuration. - Collaborate across SSS, , and deployment configurations to ensure stack lifecycle workflows remain reliable, observable, and resilient. - Improve operational efficiency by reducing deployment complexity (e.g., aiming for single PR regional SSS deployment) and contributing to the Stack Config Reconciliation project. - You may not meet every requirement, and that’s okay. If this role excites you, we’d love you to raise your hand for what could be a truly career-defining opportunity - We are seeking a Staff Backend Engineer who thrives on building production systems where correctness, scalability, and operational clarity are paramount - As a remote-first organization, you should be comfortable collaborating asynchronously across time zones and taking full ownership of the critical systems powering Grafana Cloud - We value engineers who can take ambiguous lifecycle requirements and transform them into explicit, modular solutions - You will be particularly successful in this role if you enjoy solving challenges related to stateful systems, eventual consistency, and reconciliation loops - You should be adept at breaking down complex systems work into safe, iterative increments while clearly communicating technical tradeoffs to both internal stakeholders and adjacent product teams - Our team is small and operates with a high degree of independence; you will be expected to lead major projects, coordinate across service boundaries, and help define the technical direction for our domain - Have professional experience with Golang and be willing to work across both backend service and application code - You write clean, robust, well-tested software that other engineers can understand, operate, and maintain - Have some experience with delivering projects from gathering requirements, and brainstorming ideas to shipping a product to the customer’s hands in a self-driven way - Experience participating in blameless incident response and writing high-quality post-incident reviews - You have worked on a big SaaS platform and dealt with common distributed systems problems (e.g. scalability, multi-tenancy, data isolation, HA, …) - You are willing to work across teams. Your work has to be aligned with the needs of other squads and external stakeholders. You make your plans transparent, bring stakeholders on board, and are open to feedback and suggestions - Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.) - Care deeply about developer and user experience and the quality of the products that you work on - Can take on complex challenges and break them down to achieve tight learning loops: to analyze, design, and build modular solutions, deliver MVPs, gather data and feedback, and then progress iteratively - Have experience with mentoring junior engineers in a collaborative but asynchronous environment - You have at least 1 year of fully remote work experience - Experience with TypeScript/Node.js - Experience with Jsonnet/Tanka, Terraform, Flux, Argo, or similar deployment/configuration tooling - Experience with Kubernetes control-plane patterns, operators, reconcilers, or desired-state systems - Experience working on SaaS provisioning, tenancy, regional expansion, plugin rollout, or customer lifecycle systems - Experience with incident response involving configuration drift, partial failure, or cross-service state mismatch
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring