← Back to job listings
HO
Senior Engineering Manager (Site Reliability)
Horizon3 · United States
About The Role
Join Horizon3 as a Senior Engineering Manager (Site Reliability). In this role, you will build and lead the Site Reliability function, establish incident management processes, and drive a culture of quality and excellence in engineering. You will also be responsible for recruiting and onboarding talented individuals, mentoring and developing your team, and leading horizontally with peer management and senior leaders. Enjoy a flexible work environment, growth opportunities, and an inclusive and diverse team.
- Construire et diriger la fonction de fiabilité du site chez Horizon3, y compris le recrutement et la gestion d'une équipe de 4 à 6 ingénieurs en fiabilité du site.
- Professionnaliser la gestion des incidents en définissant et en documentant les processus et pratiques d'incidents pour votre équipe SRE et pour les équipes de fonctionnalités d'application.
- Équilibrer la réponse aux incidents tout en exécutant une feuille de route d'initiatives d'observabilité et d'ingénierie de la fiabilité.
- Strong working knowledge of at least one major cloud provider (AWS, GCP, or Azure) — infrastructure, networking, managed services, IAM, cost management. AWS preferred
- Deep hands on experience with observability: application performance management, logs and traces, and golden signals and service-specific metrics
- Demonstrated experience leading hiring and growing SRE or Infrastructure teams. Experience leading or building SRE functions, including incident management processes, on-call programs, SLO/SLA definition, and operational runbooks
- Experience in selecting and deploying incident management tooling (e.g., PagerDuty, FireHydrant,etc.) and creating decision making frameworks for vendor selection
- Previous career experience as a Site Reliability Engineer. Comfortable in being hands-on while you grow and hire your team
- Ability to engage with Product and Engineering leaders to define SLOs and drive culture of ownership across teams
- Able to write clear, durable process documentation that engineering teams can adopt
- Experience in design practices like architecture decision records or RFCs in the infrastructure or platform engineering context
- Experience in providing education and training on runbook development, on-call expectations, incident investigations, and running effective postmortems
- Proven ability to to build, scale, and retain high performing engineering teams in remote or distributed environments
- Strong communication skills across technical and non-technical audiences
- Proven ability to hire, develop, and retain high-performing engineers and engineering managers in remote or distributed environments
- Proven skills in managing a backlog of strategic roadmap initiatives and a competing stream of operational needs against capacity constraints
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring