Skip to content
← Back to job listings

Staff Software Engineer (AI Reliability)

Anthropic · New York, United States

External listingfull-time24 days ago

About The Role

Join Anthropic, a leading AI safety and research company, as a Staff Software Engineer in AI Reliability. In this role, you will work closely with various teams to improve the reliability of our AI systems, specifically focusing on the serving infrastructure for our language model, Claude. You will be responsible for developing service level objectives, designing monitoring systems, leading incident response, and supporting the reliability of safeguard model serving. This position offers a comprehensive benefits package, including health insurance, paid parental leave, flexible time off, and more.

  • Develop and implement monitoring and observability systems across the token path, ensuring high availability and performance.
  • Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements.
  • Collaborate with partner teams to make the systems that deliver Claude more robust and resilient, and assist in the design and implementation of high-availability serving infrastructure.
  • Bring diverse experience -- the team's strength comes from people who've built product stacks, scaled databases, run massive distributed systems, and everything in between
  • Are curious and brave -- comfortable jumping into unfamiliar systems during an incident and helping drive resolution even when you don't have deep expertise yet
  • Think holistically about how systems compose and where the seams are
  • Care about users and feel ownership over outcomes, even for systems you don't own
  • Have strong distributed systems, infrastructure, or reliability backgrounds -- we're looking for reliability-minded software engineers and SREs
  • Can build lasting relationships across teams -- our engagement model depends on being welcomed as teammates, not outsiders with opinions
  • Have excellent communication and collaboration skills -- you'll be partnering across the entire company
  • Have been an SRE, Production Engineer, or in similar reliability-focused roles on large scale systems
  • Have experience operating large-scale model serving or training infrastructure (>1000 GPUs)
  • Have experience with one or more ML hardware accelerators (GPUs, TPUs, Trainium)
  • Understand ML-specific networking optimizations like RDMA and InfiniBand
  • Have expertise in AI-specific observability tools and frameworks
  • Have experience with chaos engineering and systematic resilience testing
  • Have contributed to open-source infrastructure or ML tooling
  • We encourage you to apply even if you do not believe you meet every single qualification
  • We require at least a Bachelor's degree in a related field or equivalent experience

This is an external listing. JobSpring does not represent or verify the employer. Report this listing