Skip to content
← Back to job listings

Senior Staff+ Software Engineer (Kubernetes Platform)

Anthropic · London, United Kingdom

External listingfull-time2 months ago

About The Role

Join Anthropic, a leading AI safety and research company, as a Senior Staff+ Software Engineer on the Kubernetes Platform team. You will be responsible for owning, operating, and extending the Kubernetes scheduler for Anthropic's accelerator fleets, scaling the Kubernetes control plane, and designing and building core cluster services. You will work closely with research, training, and inference teams to understand workload shapes and turn their requirements into platform capabilities. This role offers a competitive salary, equity packages, and comprehensive benefits.

  • Posséder, exploiter et étendre le plan de contrôle Kubernetes pour les flottes d'accélérateurs d'Anthropic, y compris les plugins et politiques de planification personnalisés.
  • Évoluer le plan de contrôle Kubernetes (apiserver, etcd, controller-manager) pour prendre en charge des clusters bien au-delà des limites typiques.
  • Concevoir, construire et exploiter des services de cluster essentiels tels que la découverte de services dont chaque charge de travail dans la flotte dépend.
  • 12+ years of relevant industry experience, including time leading large, ambiguous infrastructure projects
  • Demonstrated ability to debug complex issues across the stack, from API behavior down to node and network-level root causes
  • Proficiency in at least one systems-appropriate language (e.g., Go, Python, Rust, or C++)
  • Deep, hands-on Kubernetes experience (well beyond "user of”) into scheduler, controllers, apiserver, or operating large multi-tenant clusters
  • Strong written and verbal communication; comfort building consensus with internal stakeholders
  • A track record of designing for reliability, correctness, and clear failure semantics in systems other engineers depend on
  • Significant software engineering experience building and operating production distributed systems
  • Experience building or operating cluster schedulers or batch systems (e.g., Kueue, Volcano, Slurm, or in-house equivalents)
  • Low-level systems experience such as Linux kernel tuning, cgroups, or eBPF
  • Experience with Kubernetes internals or contributions: kube-scheduler / scheduling framework, apiserver, etcd, client-go, controller-runtime, or similar
  • We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work
  • Background scaling control planes or coordination systems (etcd, ZooKeeper, Consul, or large DNS/service-mesh deployments)
  • Familiarity with ML infrastructure: GPUs, TPUs, or Trainium; gang scheduling; topology-aware placement; collective networking such as NCCL
  • Experience with GCP and/or AWS, including GKE/EKS internals and Infrastructure as Code

This is an external listing. JobSpring does not represent or verify the employer. Report this listing