Skip to content
← Back to job listings

Senior Site Reliability Engineer

London Stock Exchange · Nottingham, United Kingdom

External listingfull-time3 months ago

About The Role

Join our Risk Intelligence division as a Senior Site Reliability Engineer. In this hands-on technical role, you will shape the foundations of reliability across new and existing platforms. Collaborate with Architecture, Engineering, Security, and Platform teams to ensure reliability is built into systems from day one. You will lead the establishment of SRE foundations for new projects, define and implement observability standards, and continuously drive reliability improvements. This position requires a proactive expert with strong leadership presence and ownership of platform reliability outcomes.

  • Establish SRE foundations for new projects, including environments, monitoring, alerting, and operational readiness.
  • Define, implement, and champion observability standards, tooling, and guidelines across metrics, logs, traces, and SLIs/SLOs.
  • Continuously drive reliability improvements across environments through incident reduction, performance tuning, and building resilient patterns.
  • Bachelor’s Degree or equivalent experience in Computer Science, Engineering, or a related field
  • Strong experience with AWS (or Azure), including services such as EKS, ECS, EC2, networking, IAM, and managed services
  • Proven experience designing and operating observability platforms, including monitoring, logging, and alerting
  • Experience with cloud cost optimization strategies and tooling
  • Experience working closely with architecture and engineering teams on system design and delivery
  • Hands-on experience with Datadog for metrics, logs, APM, and alerting
  • Strong understanding of SRE principles, including SLOs, error budgets, incident management, and reliability engineering
  • Solid understanding of cloud security principles and experience collaborating with security teams
  • Strong background in Linux systems administrations
  • 5+ years of hands-on technical experience in SRE, Platform Engineering, Infrastructure, or related roles
  • Experience supporting multi-cloud or hybrid environments
  • Exposure to Infrastructure as Code (e.g., Terraform, CloudFormation)
  • Experience in large-scale, complex, or regulated environments
  • Knowledge of vector databases and RAG architectures for building internal SRE knowledge assistants
  • Knowledge of Generative AI and LLM platforms (e.g., Claude, Amazon Bedrock)
  • Strong technical authority with the ability to influence design and operational decisions
  • Clear communicator, able to explain complex technical concepts to diverse audiences
  • Pragmatic problem-solver who balances reliability, security, cost, and delivery speed
  • Highly collaborative, comfortable working across architecture, engineering, security, and operations teams
  • Calm and methodical under pressure, especially during incidents and critical issues

This is an external listing. JobSpring does not represent or verify the employer. Report this listing