Skip to content
← Back to job listings

SRE Leader

Kontakt.io · United States

External listingfull-timeabout 1 month ago

About The Role

Join our team as an SRE Leader, where you will be responsible for the reliability, performance, and automation of our cloud-based, real-time platform. You will lead and scale the SRE team, ensuring our infrastructure meets the needs of our growing healthcare customers. Your role will involve designing and implementing self-healing systems, managing scalable cloud infrastructure, optimizing containerized environments, and driving technical strategy. You will also collaborate with various teams to align SRE initiatives with business priorities and ensure compliance with security and regulatory standards.

  • Assurer la fiabilité, la performance et l'automatisation de la plateforme cloud en temps réel, en visant un temps de disponibilité de 99,99 %.
  • Diriger et développer l'équipe SRE pour garantir que l'infrastructure reste en avance sur la demande et fonctionne efficacement.
  • Concevoir et mettre en œuvre des systèmes auto-réparateurs et tolérants aux pannes pour prévenir les défaillances avant qu'elles ne se produisent.
  • Deep expertise in cloud platforms (AWS), Kubernetes, and distributed systems
  • Strong background in monitoring, logging, and observability with Prometheus, OpenTelemetry, or similar tools
  • Deep knowledge of CI/CD automation, GitOps, and infrastructure as code (Terraform, etc.)
  • 10+ years of experience in Site Reliability Engineering or Cloud Infrastructure
  • Proven success scaling high-traffic, mission-critical platforms in SaaS, IoT, or healthcare
  • Strong understanding of network security, access management, and compliance frameworks (HIPAA, SOC 2)
  • A mature leadership approach, with the ability to drive technical strategy while growing and mentoring a high-performance SRE team
  • Hands-on experience with incident management, postmortems, and building resilient systems
  • Experience with healthcare IT, including EHR data, FHIR, and HL7 interoperability
  • Prior experience leading on-call rotations and major incident management processes
  • Expertise in real-time distributed systems, event-driven architectures, or large-scale data pipelines

This is an external listing. JobSpring does not represent or verify the employer. Report this listing