← Back to job listings
CO
Technical Program Manager (Cluster Orchestration & Applied Training)
CoreWeave · Sunnyvale, United States
About The Role
Join CoreWeave, a leading cloud provider specializing in GPU-accelerated computing for AI and machine learning. As a Staff Technical Program Manager, you will lead complex, cross-functional programs across Cluster Orchestration and Applied Training. You will partner with engineering, product, infrastructure, and research-adjacent teams to improve workload management and user interaction with the training platform. This role requires strong technical depth, excellent execution instincts, and the ability to bring structure and clarity to fast-moving infrastructure and AI platform initiatives.
- Conduire l'exécution de programmes de bout en bout pour les initiatives d'orchestration de clusters, y compris la planification des charges de travail, l'approvisionnement en libre-service, les flux de mise à niveau et de migration.
- Diriger des programmes interfonctionnels qui améliorent la manière dont l'entraînement, l'évaluation, l'apprentissage par renforcement et les charges de travail mixtes fonctionnent sur les clusters CoreWeave.
- Établir des mécanismes de programme pour la préparation au lancement, la planification du déploiement, la gestion des risques, la communication avec les parties prenantes et l'examen post-lancement.
- Demonstrated ability to define program metrics and deliver measurable outcomes in performance, reliability, scale, or operational maturity
- Strong technical fluency in Kubernetes, Slurm or comparable schedulers, distributed systems, and AI training workflows
- 8+ years of technical program management experience in cloud infrastructure, distributed systems, or AI/ML platforms
- Excellent communication skills, with experience influencing engineering, product, and executive stakeholders
- Experience leading large-scale cross-functional programs involving scheduling systems, cluster infrastructure, or ML platform capabilities
- Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
- Experience with orchestration and scheduling technologies such as Kubernetes, Slurm, Kueue, Ray, or similar systems
- Familiarity with modern AI training and evaluation workflows, including pre-training, supervised fine-tuning, reinforcement learning, and experiment or sandbox environments
- Understanding of GPU infrastructure, cluster capacity planning, multi-tenant execution, and distributed training tradeoffs
- Experience building launch processes, release governance, dependency management, and operational review mechanisms in fast-scaling environments
- Familiarity with AI developer and research tooling such as W&B, SkyPilot, or adjacent ecosystem platforms
- Wondering if you’re a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match
- This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. § 1157, or (iv) asylee under 8 U.S.C. § 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring