Skip to content
← Back to job listings

Senior Machine Learning Infrastructure Engineer (Simulation)

Waymo · Mountain View, CA, United States

External listingfull-time2 months ago

About The Role

Join Waymo, a leader in autonomous driving technology. As a Senior Machine Learning Infrastructure Engineer, you will lead the development of advanced AI/ML infrastructure for multi-billion parameter foundation models in ML accelerator-friendly simulations. You will collaborate closely with core teams, provide deep technical leadership, design and scale large distributed systems, and mentor junior engineers. Enjoy a comprehensive benefits package, including medical, dental, and vision insurance, competitive compensation, and a hybrid work model.

  • Lead the development of advanced AI/ML infrastructure for multi-billion parameter foundation models in ML accelerator-friendly simulations.
  • Design and scale large distributed systems covering the ML lifecycle, supporting planet-scale dataset generation and model training.
  • Provide deep technical leadership on large-scale ML model architectures, especially for autonomous vehicle models, and mentor junior engineers.
  • BS in Computer Science, Robotics, similar technical field of study, or equivalent practical experience
  • 5+ years of professional software engineering experience, with at least 3 years in machine learning infrastructure such as developing, scaling, training, deploying, and optimizing large-scale machine learning systems from data to model
  • 10+ years of professional software engineering experience, with at least 5 years in machine learning infrastructure such as developing, designing, scaling, training, deploying, and optimizing large-scale machine learning systems from data to model
  • Solid experience in the development and optimization of machine learning infrastructure tools like DeepSpeed, PyTorch, TensorFlow, or similar frameworks
  • Strong expertise in distributed training techniques, including gradient sharding and optimization strategies for scaling large models across ML accelerator profiling tools to uncover performance bottlenecks
  • Deep understanding of state-of-the-art machine learning models such as auto-regressive transformers and familiarity with custom-kernels for diverse h/w compute based efficiency
  • Excellent communication skills, both verbal and written, with the ability to translate complex technical concepts for a broad audience
  • Practical familiarity in Autonomous Driving, Simulations, and ML accelerators is a plus

This is an external listing. JobSpring does not represent or verify the employer. Report this listing