Skip to content
← Back to job listings

Staff Software Engineer (Cluster Orch (SUNK))

CoreWeave · Sunnyvale, United States

External listingfull-time27 days ago

About The Role

Join CoreWeave, a leading cloud provider for AI workloads. As a Staff Software Engineer, you will play a key role in advancing our orchestration platform, ensuring seamless and efficient workload management across massive GPU clusters. You will define architectural direction, mentor senior engineers, and drive cross-organizational initiatives. This is an opportunity to shape the future of AI infrastructure and make a significant impact in the industry.

  • Contribuer à l'avancement de la plateforme d'orchestration de CoreWeave, y compris SUNK (Slurm sur Kubernetes) et au-delà.
  • Définir la direction architecturale, posséder des parties critiques de la plateforme d'orchestration et conduire des initiatives inter-organisationnelles.
  • Mentorer les ingénieurs seniors, établir des meilleures pratiques au sein de l'organisation en matière de fiabilité et d'observabilité.
  • Advanced proficiency in Go and distributed systems design
  • 8–12 years of professional software engineering experience
  • Ability to mentor senior engineers and elevate organizational standards
  • Experience setting technical direction and influencing cross-team architecture
  • Deep expertise in Slurm/Kubernetes internals and cloud-native development
  • Proven track record designing and operating large-scale distributed systems in production
  • Experience with distributed workloads, GPU-based applications, or ML pipelines
  • Familiarity with orchestration and workflow technologies such as Ray, Kubeflow, Kueue, Istio, Knative, or Argo Workflows
  • Knowledge of scheduling concepts like quota enforcement, pre-emption, and scaling strategies
  • Experience with AI infrastructure and workloads (ML training, inference, or HPC)
  • You’re curious about orchestration beyond SUNK and how to evolve it for next-generation AI
  • Exposure to reliability practices including SLOs, alarms, and post-incident reviews
  • You’re an expert at mentorship, architecture, and operational excellence across teams
  • You love defining long-term architecture for systems at global scale
  • You thrive on solving problems that balance cost, performance, and reliability in high-demand environments

This is an external listing. JobSpring does not represent or verify the employer. Report this listing