Skip to content
← Back to job listings

Senior Staff Software Engineer (Node Infrastructure)

Anthropic · London, United Kingdom

External listingfull-time24 days ago

About The Role

Join Anthropic, a leading AI safety and research company. As a Senior Staff Software Engineer in the Node Infrastructure team, you will own the technical strategy and roadmap for node lifecycle management, drive cross-team initiatives, and design and operate systems for hardware health and repair. You will work closely with cloud providers and internal teams to shape long-term infrastructure strategy and support the growth of engineers through mentorship and coaching. Enjoy comprehensive benefits, a competitive salary, and the opportunity to make a significant impact in the field of AI.

  • Own the technical strategy and roadmap for node lifecycle management, including ingestion, bring-up, health checking, and automated repair.
  • Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families, ensuring the hardest problems get solved.
  • Establish and evolve operational excellence practices, including incident response, postmortem culture, and on-call, while supporting the growth of engineers through technical mentorship and coaching.
  • Hands-on experience with machine learning accelerators (GPUs, TPUs, or Trainium)
  • Deep expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, AWS/GCP/Azure)
  • Track record of leading complex, multi-quarter technical initiatives that span multiple teams or systems
  • Strong proficiency in at least one systems language (e.g., Rust, Go, or Python), IaC proficiency with Terraform
  • Ability to build alignment across senior stakeholders and communicate effectively at all levels
  • Experience managing large scale compute infrastructure at hyperscale (10K+ nodes), including capacity management and efficiency
  • Depth in one or more of: Kubernetes internals (scheduler, autoscaler, kubelet, Karpenter), cluster orchestration systems (Mesos, Borg-like), or node provisioning pipelines
  • Low-level systems experience: kernel, virtualization, device drivers, firmware, or hardware health/diagnostics daemons
  • Familiarity with high-performance networking (EFA, RDMA, InfiniBand) for distributed ML workloads
  • Contributions to relevant open-source projects (Kubernetes, Linux kernel, container runtimes, etc.)
  • Demonstrated ownership of production reliability for high-throughput, latency-sensitive systems
  • Skill in quickly understanding systems design tradeoffs and keeping track of rapidly evolving software systems
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
  • Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
  • We encourage you to apply even if you do not believe you meet every single qualification
  • 12+ years of software engineering experience, including time as a technical lead setting direction for a team

This is an external listing. JobSpring does not represent or verify the employer. Report this listing