Skip to content
← Back to job listings

Research Engineer (Discovery)

Anthropic · San Francisco, United States

External listingfull-time24 days ago

About The Role

Join our team as a Research Engineer, where you will play a crucial role in advancing scientific AGI. You will design and implement large-scale infrastructure systems, identify and resolve infrastructure bottlenecks, and develop robust evaluation frameworks. You will also build scalable VM/sandboxing/container architectures, collaborate on translating experimental requirements into production-ready infrastructure, and optimize training and inference pipelines. This position offers comprehensive benefits, a competitive salary, and the opportunity to work on cutting-edge AI research.

  • Identifying and addressing key infrastructure blockers on the path to scientific AGI.
  • Designing and implementing large-scale infrastructure systems to support AI scientist training, evaluation, and deployment.
  • Developing robust and reliable evaluation frameworks for measuring progress towards scientific AGI.
  • Have 6+ years of highly relevant experience in infrastructure engineering with demonstrated expertise in large-scale distributed systems
  • Strong candidates should have familiarity with performance optimization, distributed systems, vm/sandboxing/container deployment, and large scale data pipelines
  • Familiarity with language model training, evaluation, and inference is highly encouraged
  • Possess deep knowledge of performance optimization techniques and system architectures for high-throughput ML workloads
  • Have experience collaborating with other researchers to scale experimental ideas
  • Excel at diagnosing and resolving complex infrastructure challenges in production environments
  • Are a strong communicator and enjoy working collaboratively
  • Can work effectively across the full ML stack from data pipelines to performance optimization
  • Have experience with containerization technologies (Docker, Kubernetes) and orchestration at scale
  • Have proven track record of building large-scale data pipelines and distributed storage systems
  • Experience with language model training infrastructure and distributed ML frameworks (PyTorch, JAX, etc.)
  • Background in building infrastructure for AI research labs or large-scale ML organizations
  • Experience with cloud platforms (AWS, GCP) at enterprise scale
  • Knowledge of GPU/TPU architectures and language model inference optimization
  • Familiarity with VM and container orchestration
  • Experience with workflow orchestration tools and experiment management systems
  • History working with large scale reinforcement learning
  • Comfort with large scale data pipelines (Beam, Spark, Dask, …)
  • Education requirements: We require at least a Bachelor's degree in a related field or equivalent experience

This is an external listing. JobSpring does not represent or verify the employer. Report this listing