Lead Site Reliability Engineer
London Stock Exchange · Nottingham, United Kingdom
About The Role
Join our Markets and Risk Intelligence division as a Lead Site Reliability Engineer. In this senior hands-on technical role, you will shape the foundations of reliability across new and existing platforms. Collaborate with Architecture, Engineering, Security, and Platform teams to ensure reliability is built into systems from day one. You will establish SRE foundations for new projects, define and implement observability standards, and continuously drive reliability improvements. This position requires a highly proactive expert with strong leadership presence and ownership of platform reliability outcomes.
- Establish SRE foundations for new projects, ensuring operational readiness from day one.
- Collaborate with Architecture and Engineering teams to embed reliability, scalability, security, and observability into system design.
- Lead the design and evolution of monitoring and alerting solutions that improve visibility, reduce toil, and strengthen system health.
- This position requires a highly proactive, hard-working expert with strong leadership presence and ownership of platform reliability outcomes
- We are looking for a person who is passionate about reliability engineering and who bring a continuous improvement approach to everything they do!
- Bachelor’s Degree in Computer Science or related field
- Strong understanding of SRE principles, including SLOs, error budgets, incident management, and reliability engineering
- Proven experience designing and operating observability platforms, including monitoring, logging, and alerting
- Hands-on experience with Datadog for metrics, logs, APM, and alerting
- Solid understanding of cloud security principles and experience collaborating with security teams
- 10+ years of hands-on technical experience in SRE, Platform Engineering, Infrastructure, or related roles
- Experience with cloud cost optimisation strategies and tooling
- Strong experience with AWS, including services such as EKS, ECS, EC2, networking, IAM, and managed services
- Hands-on experience integrating AI with observability stacks (Prometheus, Grafana, ELK, OpenTelemetry) for proactive issue detection
- Deep hands-on experience with Kubernetes and containerised platforms
- Experience working closely with architecture and engineering teams on system design and delivery
- Strong background in Linux systems administrations
- Experience or working knowledge of Microsoft Azure
- Experience supporting multi-cloud or hybrid environments
- Exposure to Infrastructure as Code (e.g., Terraform, CloudFormation)
- Experience in large-scale, complex, or regulated environments
- Knowledge of vector databases and RAG architectures for building internal SRE knowledge assistants
- Knowledge of Generative AI and LLM platforms (e.g., Claude, Amazon Bedrock)
- Strong technical authority with the ability to influence design and operational decisions
- Calm and methodical under pressure, especially during incidents and critical issues
- Highly collaborative, comfortable working across architecture, engineering, security, and operations teams
- Pragmatic problem-solver who balances reliability, security, cost, and delivery speed
- Clear communicator, able to explain complex technical concepts to diverse audiences
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring