Skip to content
← Back to job listings

Lead Site Reliability Engineer

Movable Ink · New York, United States

External listingfull-time16 days ago

About The Role

Join our team as a Lead Site Reliability Engineer, where you will combine technical expertise with strategic leadership. You will design and evolve major systems within our multi-cloud, multi-region content serving platform, driving reliability initiatives and defining the technical strategy for scaling our platform. You will also own the design and evolution of core platform applications, establish capacity planning and performance management frameworks, and lead cross-functional reliability initiatives.

  • Concevoir et faire évoluer des systèmes majeurs au sein de notre plateforme de diffusion de contenu multi-cloud.
  • Définir et mettre en œuvre la stratégie d'automatisation pour les outils d'infrastructure, en établissant des normes qui minimisent le travail manuel.
  • Diriger des initiatives de fiabilité interfonctionnelles avec les équipes SRE et d'ingénierie des services, en influençant les décisions architecturales.
  • Proven track record in Site Reliability or Software Engineering, designing, building, and owning scalable, resilient services with a focus on long-term reliability strategy
  • Experience architecting and leading large-scale observability platforms, including defining observability standards and SLO frameworks. We use Prometheus and Thanos with Grafana Alloy, Loki and Tempo
  • Designing and owning automation strategies to manage services at scale, with expertise in establishing performance analysis frameworks and mentoring others on diagnostics and resolution
  • Deep expertise in architecting and operating complex distributed systems such as Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB/Cassandra, with the ability to guide teams through distributed system challenges
  • Deep, hands-on experience (6+ years) in Site Reliability or Software Engineering, specifically leading and shaping multi-cloud architecture and strategy (AWS and GCP)
  • Advanced Linux systems expertise, with the ability to diagnose complex system-level issues and mentor others on performance tuning and troubleshooting
  • Experience leading on-call excellence, including driving improvements to monitoring and alerting strategies, automating runbooks and mentoring team members on incident response best practices. Every member of the SRE team does a week long on-call rotation
  • Advanced Kubernetes expertise, including cluster architecture design, multi-tenancy strategies, and guiding teams on container orchestration best practices. We use EKS and GKE
  • Proficiency in multiple programming languages with the ability to design and review code that meets reliability standards. We use NodeJS, Golang, Ruby, Python and shell scripting
  • Expert-level proficiency with infrastructure as code, including defining IaC standards and patterns across teams. We use Terraform and Chef
  • If you’re excited about the role but don’t meet all of the abovementioned qualifications, we encourage you to apply

This is an external listing. JobSpring does not represent or verify the employer. Report this listing