Senior Software Engineer (Applied Training)
CoreWeave · Sunnyvale, United States
About The Role
Join CoreWeave as a Senior Software Engineer (Applied Training) and be an early member of a small team responsible for developing our Kubernetes-native research cluster platform. You will contribute to the roadmap for Applied Training, work closely with customers and other teams, and design and build a complete research cluster experience. You will also own the Python SDK for sandbox infrastructure, write documentation for popular OSS training frameworks, and work directly with infrastructure teams and customers. The ideal candidate has 8-12+ years of experience building distributed systems, ML infrastructure, or developer platforms, and has real Kubernetes experience.
- Contribuer à la feuille de route pour la formation appliquée, en identifiant les éléments clés qui débloquent de nouvelles charges de travail.
- Travailler directement avec les clients et d'autres équipes pour concevoir et construire une expérience complète de cluster de recherche.
- Écrire de la documentation pour l'exécution de frameworks de formation OSS populaires sur CoreWeave afin de débloquer les clients.
- Familiarity with training: how distributed jobs get scheduled, how ranks initialize, what breaks at scale
- You understand what makes researchers productive: code distribution matters, fast iteration cycles matter, workflows that don't require becoming infrastructure experts matter
- Strong communicator who can work with customers and translate researcher complaints into system designs
- You've shipped infrastructure that other people rely on daily — not prototypes, production systems
- 8–12+ years building distributed systems, ML infrastructure, or developer platforms
- Real Kubernetes experience: custom controllers, operators, scheduling, CRDs, workload orchestration at scale — not just deploying things to Kubernetes or cluster administration
- Familiarity with agentic AI: RL training with rollouts, agent evaluation, sandbox isolation for running untrusted code
- Experience building internal ML platforms or research clusters at a company doing large-scale training
- Experience with container runtimes, isolation (gVisor, Kata), or serverless platforms
- Background with Slurm, Ray, or similar workload orchestration, and opinions on where they fall short
- OSS contributions to Kubernetes SIGs, Ray, PyTorch, or similar
- Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring