Skip to content
← Back to job listings

Staff Software Engineer (Databases SRE)

Grafana Labs · Ireland

External listingfull-time28 days ago

About The Role

Join Grafana Labs as a Staff Software Engineer - SRE, where you'll focus on increasing the reliability of our Cloud databases. You'll partner closely with product engineering squads, own production reliability for high-SLA and complex customer environments, and design and implement automation to scale our reliability practices. This is a fully remote position with a range of benefits including 30 days of paid vacation, healthcare coverage, retirement planning, and professional development opportunities.

  • Collaborer étroitement avec les équipes d'ingénierie produit pour assurer la fiabilité de nos bases de données Cloud.
  • Concevoir et mettre en œuvre des automatisations pour améliorer nos pratiques de fiabilité et réduire les incidents.
  • Servir de point d'escalade principal et de responsable des incidents pour les environnements clients complexes.
  • Experience with one or more programming languages (e.g. Go, Python, Java, etc)
  • Ability to reason about performance, scaling, and failure modes
  • Ability to partner deeply with product engineering teams
  • Comfortable working within an engineering team where individuals are encouraged to have a strong sense of autonomy and self-direction
  • Excellent problem-solving and troubleshooting skills
  • Strong experience with technical leadership, leading a team through projects, mentoring other engineers on the team and serving as a force-multiplier
  • We highly value those who are intellectually curious, who default to transparency, possess a high bias towards action, and who are also kind (this is important!)
  • Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents)
  • 8+ years engineering experience, 4+ in SRE/CRE/production engineering. Strong preference for those with formal customer reliability engineering experience
  • Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.)
  • Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling
  • Strong experience designing and implementing SLOs
  • Experience operating multi-tenant systems in production

This is an external listing. JobSpring does not represent or verify the employer. Report this listing