← Back to job listings
CE
Senior Manager of Infrastructure Platform Engineering
Crusoe Energy Systems · Sunnyvale, United States
About The Role
Join our team as a Senior Manager of Infrastructure Platform Engineering. In this hands-on management role, you will lead a team of infrastructure software engineers and set the technical direction for our core systems that turn large-scale compute infrastructure into reliable, secure, and efficiently allocatable capacity. You will partner closely with adjacent infrastructure, production engineering, and security teams to ensure the reliability and usability of our platform. This role is essential to the customer experience on our platform and directly underpins the business.
- Lead a team responsible for building core systems that turn large-scale compute infrastructure into reliable, secure, and efficiently allocatable capacity.
- Set technical direction across the platform, driving the design of secure, well-instrumented platform systems, and establishing engineering standards for infrastructure software development.
- Hire, mentor, and grow a team of infrastructure software engineers, building a high-performing organization from a strong foundation.
- Deep expertise in large-scale infrastructure platforms — building services that pool, allocate, and reconcile compute resources at scale
- Track record of hiring strong infrastructure engineers and helping them grow into more senior roles
- Strong background with Kubernetes and cloud platforms (GCP, AWS, or Azure) — orchestration, automation, and operating distributed systems in production
- Experience with efficiency, capacity, or performance engineering — characterizing system behavior, identifying bottlenecks, and driving measurable improvements in utilization or availability
- 10+ years of experience in infrastructure or systems software development, with at least 3+ years in an engineering leadership role
- Comfortable operating in a fast-moving environment where the path isn't fully paved — willing to drive ambiguity to clarity
- Experience with distributed state management and control systems — modeling resource and system lifecycle, reconciling desired vs. actual state, and handling failure gracefully across a large fleet
- A player-coach approach to management: hands-on enough to make technical calls, structured enough to grow a team and ship through them
- Experience operating Kubernetes on bare-metal infrastructure as well as on managed cloud services (GKE, EKS, AKS)
- Familiarity with the operational challenges of GPU clusters, AI training, and inference workloads
- Working knowledge of platform security and trust concepts — secure boot, measured boot, TPMs, and hardware attestation
- Experience with capacity forecasting, demand modeling, or allocation optimization at scale
- Hands-on background with telemetry and observability platforms at scale (Prometheus, OpenTelemetry, Grafana)
- Prior experience building infrastructure platforms at hyperscalers or cloud providers where internal engineers are the primary customer
- Familiarity with hardware-software co-design — understanding how platform choices affect physical infrastructure utilization
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring