Skip to content
← Back to job listings

Site Reliability Engineer

Reward Gateway · London, United Kingdom

External listingfull-time16 days ago

About The Role

Join our team as a Site Reliability Engineer and help us transform our operational workloads to an SRE approach. You will work closely with our Product Engineering teams, implement a new standard of observability, and advocate for availability, reliability, and uptime. You will also take part in SRE Incident Management processes, act as a key Incident Commander, and ensure cost efficiency for our platforms and services. Enjoy a range of benefits including life assurance, pension, unlimited free books, flexible working, and more.

  • Integrating tightly with Product Engineering teams and following SRE practices to maintain high standards of compliance.
  • Implementing a new standard of observability utilizing SLI/SLO/Error Budgets and continually evolving observability platforms.
  • Actively participating in SRE Incident Management processes and acting as a key Incident Commander within the Incident Management process.
  • Managing services using SLI/SLO & Error Budgets
  • Good experience in DevOps or SRE, with a keen interest to learn and grow as a Site Reliability Engineer
  • An ability to learn new tools and processes quickly and impart that knowledge
  • Ability to work under pressure and be highly reliable
  • Adaptability and flexibility to change in a fast-moving environment
  • Some understanding of SQL, PHP, Kubernetes, CI/CD advantageous
  • Ability to work both independently and as part of a team
  • Experience in HA environments
  • Experience with AWS or other cloud providers
  • Good SRE skills with a good understanding of SRE practices
  • Automation skills through Terraform, Python, Bash or similar
  • Observability product experience (eg Datadog)

This is an external listing. JobSpring does not represent or verify the employer. Report this listing