Skip to content
← Back to job listings

Senior Staff Software Engineer (Node Infrastructure)

Anthropic · New York, United States

External listingfull-time20 days ago

About The Role

Join Anthropic, a leading AI safety and research company, as a Senior Staff Software Engineer. In this role, you will own the technical strategy and roadmap for node lifecycle management, drive cross-team initiatives to build and scale AI clusters, and design and operate systems for automatic hardware remediation. You will also define infrastructure architecture, establish operational excellence practices, and support the growth of engineers through mentorship and coaching.

  • Own the technical strategy and roadmap for node lifecycle management, including ingestion, bring-up, health checking, and automated repair.
  • Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families, ensuring the infrastructure architecture is defined and evolved.
  • Establish and evolve operational excellence practices, including incident response, postmortem culture, and on-call support, while supporting the growth of engineers through mentorship and coaching.
  • We encourage you to apply even if you do not believe you meet every single qualification. We urge you not to exclude yourself prematurely and to submit an application if you're interested in this work
  • 12+ years of software engineering experience, including time as a technical lead setting direction for a team
  • Strong proficiency in at least one systems language (e.g., Rust, Go, or Python), IaC proficiency with Terraform
  • Hands-on experience with machine learning accelerators (GPUs, TPUs, or Trainium)
  • Deep expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, AWS/GCP/Azure)
  • Track record of leading complex, multi-quarter technical initiatives that span multiple teams or systems
  • Ability to build alignment across senior stakeholders and communicate effectively at all levels
  • Experience managing large scale compute infrastructure at hyperscale (10K+ nodes), including capacity management and efficiency
  • Skill in quickly understanding systems design tradeoffs and keeping track of rapidly evolving software systems
  • Low-level systems experience: kernel, virtualization, device drivers, firmware, or hardware health/diagnostics daemons
  • Contributions to relevant open-source projects (Kubernetes, Linux kernel, container runtimes, etc.)
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
  • Familiarity with high-performance networking (EFA, RDMA, InfiniBand) for distributed ML workloads
  • Demonstrated ownership of production reliability for high-throughput, latency-sensitive systems
  • Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
  • Depth in one or more of: Kubernetes internals (scheduler, autoscaler, kubelet, Karpenter), cluster orchestration systems (Mesos, Borg-like), or node provisioning pipelines

This is an external listing. JobSpring does not represent or verify the employer. Report this listing