← Back to job listings
DA
Site Reliability Engineer
Darktrace · Cambridge, United Kingdom
About The Role
Join our team as a Site Reliability Engineer (SRE) and play a key role in shaping the future of our platform reliability strategy. In this position, you will act as the go-to authority in your area of expertise, working across teams to embed best practices, solve complex reliability challenges, and improve system resilience at scale. You will focus on a core domain of expertise, such as observability, performance engineering, data infrastructure reliability, security-focused SRE, or network reliability, while influencing reliability standards across the wider engineering organization.
- Act as the subject matter expert in your chosen reliability domain, defining and implementing standards, frameworks, and best practices across SRE, Platform Engineering, and DevSecOps.
- Design and implement solutions to complex, cross-cutting reliability challenges, building tooling, automation, and frameworks to improve system resilience and scalability.
- Play a key role in incident response, particularly within your specialism, contributing to on-call rotations and continuous improvement of operational processes.
- Security-focused SRE (hardening, compliance automation, secrets management)
- Strong communication skills, with the ability to explain complex technical concepts clearly
- Observability & monitoring (metrics, logging, distributed tracing)
- Experience with cloud platforms (AWS, GCP, Azure) and Kubernetes
- Self-driven with the ability to identify and prioritise high-impact work independently
- Data infrastructure reliability (databases, streaming, pipelines)
- Proven experience in Site Reliability Engineering, DevOps, or infrastructure engineering
- Network reliability & traffic management
- Performance engineering & capacity planning
- Deep expertise in at least one of the following areas:
- Strong programming skills (e.g. Go, Python, or similar)
- Experience building internal developer platforms or tooling
- Contributions to open-source, technical blogs, or public speaking
- Experience working in regulated environments
- Relevant certifications in your specialist domain
- Familiarity with SLO frameworks and error budget management
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring