Skip to content
← Back to job listings

Site Reliability Engineer (SRE) - AI GPU Clusters

Scaleway · Paris, Ile-de-France, France

External listingfull-timeabout 1 month ago

About The Role

Join Scaleway, a leading European cloud computing company, as a Site Reliability Engineer (SRE) focused on AI GPU clusters. In this role, you will build and maintain reliable, observable, and secure infrastructure to ensure optimal service availability for customers worldwide. You will work in a collaborative and international environment, contributing to the design, maintenance, and scaling of core systems and observability tools.

  • Construire et maintenir une infrastructure AI fiable, observable et sécurisée pour garantir la disponibilité optimale des services.
  • Participer à la rotation d'astreinte pour gérer les incidents et assurer la continuité du service.
  • Mettre en œuvre et maintenir des solutions d'observabilité pour surveiller l'infrastructure AI et la santé des applications.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing