← Back to job listings
AN
Staff and Senior Software Engineer (Inference)
Anthropic · New York, United States
About The Role
Join our Inference team as a Staff or Senior Software Engineer. You will be responsible for building and maintaining the critical systems that serve our AI model, Claude, to millions of users worldwide. Your work will involve designing and implementing intelligent request routing, load balancing, and traffic management systems, as well as maximizing compute efficiency and supporting inference for new model architectures. You will have the opportunity to work with cutting-edge AI technology and make a significant impact on the future of AI.
- Design, build, and maintain the distributed systems that serve Claude to millions of users worldwide, including intelligent request routing and load balancing.
- Maximize compute efficiency across the fleet by autoscaling and orchestrating production, research, and experimental workloads, and build production-grade deployment pipelines.
- Integrate new AI accelerator platforms and support inference for new model architectures, while analyzing observability data to tune performance based on real-world production workloads.
- Desire to learn more about machine learning systems and infrastructure
- Thrive in environments where technical excellence directly drives both business results and research breakthroughs
- Results-oriented, with a bias towards flexibility and impact
- Significant software engineering experience, particularly with distributed systems
- Enjoy pair programming (we love to pair!)
- Care about the societal impacts of your work
- Willingness to pick up slack, even if it goes outside your job description
- Experience with load balancing, request routing, or traffic management systems
- Experience implementing and deploying machine learning systems at scale
- Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
- Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
- Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure)
- Familiarity with LLM inference optimization, batching, and caching strategies
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
- Proficiency in Python or Rust
- Experience with high-performance, large-scale distributed systems
- We encourage you to apply even if you do not believe you meet every single qualification
- Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring