Skip to content
← Back to job listings

Research Engineer (Pretraining Scaling)

Anthropic · London, United Kingdom

External listingfull-time22 days ago

About The Role

Join Anthropic's ML Performance and Scaling team as a Research Engineer. In this role, you will ensure the reliable, efficient, and scalable training of our frontier models. You will work across our entire production training stack, including performance optimization, hardware debugging, experimental design, and launch coordination. This is a demanding, high-impact position that requires deep technical expertise and a passion for large-scale ML systems. You will have the opportunity to gain hands-on experience with some of the largest training runs in the industry and work alongside world-class researchers and engineers.

  • Assurer le bon fonctionnement, l'efficacité et l'évolutivité des modèles de préformation de l'entreprise.
  • Déboguer et résoudre des problèmes complexes à travers l'ensemble de la pile technologique, de l'infrastructure matérielle aux dynamiques d'entraînement.
  • Concevoir et exécuter des expériences pour améliorer l'efficacité de l'entraînement, réduire le temps d'étape, augmenter le temps de disponibilité et améliorer les performances du modèle.
  • Experience with production ML systems, observability tools, or evaluation infrastructure
  • Care about the societal impacts of AI and responsible scaling
  • Genuinely enjoy both research and engineering work—you'd describe your ideal split as roughly 50/50 rather than heavily weighted toward one or the other
  • Have hands-on experience training large language models, or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems
  • Contributed to open-source LLM frameworks (e.g., open_lm, llm-foundry, mesh-transformer-jax)
  • Published research on model training, scaling laws, or ML systems
  • Background as a systems engineer, quant, or in other roles requiring both technical depth and operational excellence
  • Previous experience training LLM’s or working extensively with JAX/TPU, PyTorch, or other ML frameworks at scale
  • Are excited about being on-call for production systems, working long days during launches, and solving hard problems under pressure
  • Thrive when working on whatever is most impactful, even if that changes day-to-day based on what the production model needs
  • Excel at debugging complex, ambiguous problems across multiple layers of the stack
  • Are passionate about the work itself and want to refine your craft as a research engineer
  • We require at least a Bachelor's degree in a related field or equivalent experience
  • Communicate clearly and collaborate effectively, especially when coordinating across time zones or during high-stress incidents
  • We encourage you to apply even if you do not believe you meet every single qualification

This is an external listing. JobSpring does not represent or verify the employer. Report this listing