← Back to job listings
CL
Engineering Manager (Production Orchestration)
Cockroach Labs · New York, United States
About The Role
Join Cockroach Labs as an Engineering Manager for the Production Orchestration team. You will lead the team responsible for the reliability, availability, and scalability of CockroachDB in production. Your role will involve driving operational excellence, automation, and foundational architecture, as well as coaching and developing your engineers. You will collaborate with engineering and product leadership to shape the roadmap for CockroachDB's operational capabilities and future products.
- Lead the Production Orchestration team, focusing on the reliability, availability, and scalability of CockroachDB in production.
- Drive automation and tooling to reduce operational toil by building systems that improve observability and scale the fleet.
- Coach and develop your engineers, providing direct, constructive feedback, and guiding personal development and career growth.
- Experience with performance management, understanding the importance of building an effective team that can function independently while collaborating and supporting each other
- Experience working on complex technical products with exposure to distributed systems, cloud infrastructure, container orchestration, or large-scale fleet management
- Comfort with programming languages like Go and Python. We use Go, but if you don't know it, you'll learn while you're here
- Solid systems architecture knowledge and an understanding of how a variety of teams' interactions may impact operational reliability
- Partnered across departments, ensuring coordination with internal teams and external partner teams across time zones
- Experience leading global operations and/or incident management and response
- A passion for building relationships and a deep sense of responsibility for the welfare of the engineering team you manage, including their professional development and growth. We're looking for managers that want to empower their team to achieve their professional and personal goals
- A strong SRE or Production Engineering background. You understand the principles of reliability engineering, SLOs/SLAs, error budgets, and the engineering approach to operations
- Experience supporting workloads across multiple cloud providers (GCP, AWS, Azure)
- Grown or managed teams that coordinate across multiple time zones
- Leveraged, or even better built, observability tooling for your team and the rest of your org
- Experience applying AI/ML to operational workflows (e.g., intelligent alerting, automated remediation, capacity forecasting)
- Familiarity with CockroachDB or distributed SQL databases
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring