← Back to job listings
AN
Research Engineer (Discovery)
Anthropic · San Francisco, United States
About The Role
Join our team as a Research Engineer, where you will play a crucial role in advancing scientific AGI. You will design and implement large-scale infrastructure systems, identify and resolve infrastructure bottlenecks, and develop robust evaluation frameworks. You will also build scalable VM/sandboxing/container architectures, collaborate on translating experimental requirements into production-ready infrastructure, and optimize training and inference pipelines. This position offers comprehensive benefits, a competitive salary, and the opportunity to work on cutting-edge AI research.
- Identifying and addressing key infrastructure blockers on the path to scientific AGI.
- Designing and implementing large-scale infrastructure systems to support AI scientist training, evaluation, and deployment.
- Developing robust and reliable evaluation frameworks for measuring progress towards scientific AGI.
- Have 6+ years of highly relevant experience in infrastructure engineering with demonstrated expertise in large-scale distributed systems
- Strong candidates should have familiarity with performance optimization, distributed systems, vm/sandboxing/container deployment, and large scale data pipelines
- Familiarity with language model training, evaluation, and inference is highly encouraged
- Possess deep knowledge of performance optimization techniques and system architectures for high-throughput ML workloads
- Have experience collaborating with other researchers to scale experimental ideas
- Excel at diagnosing and resolving complex infrastructure challenges in production environments
- Are a strong communicator and enjoy working collaboratively
- Can work effectively across the full ML stack from data pipelines to performance optimization
- Have experience with containerization technologies (Docker, Kubernetes) and orchestration at scale
- Have proven track record of building large-scale data pipelines and distributed storage systems
- Experience with language model training infrastructure and distributed ML frameworks (PyTorch, JAX, etc.)
- Background in building infrastructure for AI research labs or large-scale ML organizations
- Experience with cloud platforms (AWS, GCP) at enterprise scale
- Knowledge of GPU/TPU architectures and language model inference optimization
- Familiarity with VM and container orchestration
- Experience with workflow orchestration tools and experiment management systems
- History working with large scale reinforcement learning
- Comfort with large scale data pipelines (Beam, Spark, Dask, …)
- Education requirements: We require at least a Bachelor's degree in a related field or equivalent experience
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring