Skip to content
← Back to job listings

Staff Software Engineer (Databases SRE)

Grafana Labs · Spain

External listingfull-timeabout 1 month ago

About The Role

Join Grafana Labs as a Staff Software Engineer - SRE, where you'll play a crucial role in enhancing the reliability of our Cloud databases. You'll work closely with product engineering squads, own production reliability for high-SLA and complex customer environments, and design and implement automation to scale our reliability practices. This is a fully remote position with a range of benefits, including 30 days of paid vacation, healthcare coverage, retirement planning, and professional development opportunities.

  • Collaborer étroitement avec les équipes d'ingénierie produit pour assurer la fiabilité de nos bases de données Cloud.
  • Concevoir et mettre en œuvre des automatisations pour améliorer nos pratiques de fiabilité et réduire les incidents.
  • Définir et faire évoluer les SLOs par client et les modèles de fiabilité, en réduisant proactivement l'épuisement des SLOs.
  • You may not meet every requirement, and that’s okay. If this role excites you, we’d love you to raise your hand for what could be a truly career-defining opportunity
  • Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling
  • Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.)
  • Experience with one or more programming languages (e.g. Go, Python, Java, etc)
  • We highly value those who are intellectually curious, who default to transparency, possess a high bias towards action, and who are also kind (this is important!)
  • Ability to partner deeply with product engineering teams
  • Excellent problem-solving and troubleshooting skills
  • 8+ years engineering experience, 4+ in SRE/CRE/production engineering. Strong preference for those with formal customer reliability engineering experience
  • Strong experience designing and implementing SLOs
  • Ability to reason about performance, scaling, and failure modes
  • Comfortable working within an engineering team where individuals are encouraged to have a strong sense of autonomy and self-direction
  • Experience operating multi-tenant systems in production
  • Strong experience with technical leadership, leading a team through projects, mentoring other engineers on the team and serving as a force-multiplier
  • Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents)

This is an external listing. JobSpring does not represent or verify the employer. Report this listing