Skip to content
← Back to job listings

Staff Cloud Site Reliability Engineer (AI/ML Platform & GPU Compute)

Wayve · London, United Kingdom

External listingfull-time22 days ago

About The Role

Join Wayve, a pioneering company in the field of AI and self-driving technology. As a Staff Cloud Site Reliability Engineer, you will play a crucial role in shaping the reliability of large-scale AI systems and GPU compute infrastructure. You will be responsible for building and scaling the reliability foundations of our AI cloud platform, defining frameworks and operational standards, and ensuring the resilience and performance of our cloud infrastructure. This is a unique opportunity to be a founding member of our Cloud SRE team and make a significant impact in the field of AI.

  • Assurer la fiabilité, la disponibilité et la performance de la plateforme de développement de modèles et des environnements de calcul GPU.
  • Participer à une rotation d'astreinte 24/7 en tant que première ligne de réponse pour les incidents liés au cloud et aux clusters.
  • Concevoir et exploiter des systèmes de surveillance, de journalisation, de traçage et d'alerte qui permettent une détection et une récupération rapides.
  • We understand that everyone has a unique set of skills and experiences and that not everyone will meet all of the requirements listed above. If you’re passionate about self-driving cars and think you have what it takes to make a positive impact on the world, we encourage you to apply
  • Strong Kubernetes experience, including operating production clusters
  • Experience running model training or inference pipelines in production (MLOps)
  • Experience working with large compute clusters; exposure to AI/ML training or inference workloads strongly preferred
  • Strong Linux fundamentals and proficiency in at least one scripting or systems language (e.g. Python, Go, C++) with a bias toward automation
  • Hands-on experience running production workloads in AWS, GCP, or Azure
  • Experience operating complex distributed systems in production, ideally including compute-heavy or high-performance workloads
  • Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large-scale cloud systems
  • Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements
  • Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry)
  • Experience operating GPU-backed environments or large-scale ML infrastructure
  • Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale
  • Experience defining and running SLOs/SLIs and building reliability programs across multiple teams
  • Familiarity with infrastructure-as-code (e.g. Terraform) and secure cloud production environments
  • Interest in helping shape and grow a Cloud SRE function, with potential to take on leadership responsibilities over time
  • Experience as an early or founding SRE hire establishing processes from scratch

This is an external listing. JobSpring does not represent or verify the employer. Report this listing