Skip to content
← Back to job listings

Senior Software Engineer (Simulation ML Infrastructure)

Waymo · San Francisco, United States

External listingfull-time22 days ago

About The Role

Join Waymo, a leader in autonomous driving technology. As a Senior Software Engineer in the Simulation ML Infrastructure team, you will lead the development of advanced AI/ML infrastructure for multi-billion parameter foundation models. You will collaborate closely with the core Waymo Realism Modeling team and work at the intersection of data engineering, model development, and simulations. This role offers competitive compensation, comprehensive benefits, and a hybrid work model.

  • Lead the development of advanced AI/ML infrastructure for multi-billion parameter foundation models in ML accelerator-friendly simulations.
  • Design and scale large distributed systems covering the ML lifecycle, supporting planet-scale dataset generation, model training, and evaluation.
  • Collaborate cross-functionally to derive performance and system-level requirements for large ML systems, translating product/business goals into measurable technical deliverables.
  • Practical familiarity in Autonomous Driving, Simulations, and ML accelerators is a plus
  • Excellent communication skills, both verbal and written, with the ability to translate complex technical concepts for a broad audience
  • Strong leadership skills with experience driving ambiguous problems end-to-end, with a willingness and independence to pick up whatever knowledge to get the job done. Passionate about building infrastructure, libraries, tools, and pipelines for engineers and scientists
  • Strong understanding of state-of-the-art machine learning models and algorithms such as autoregressive transformers and familiarity scaling large models across ML accelerator profiling tools to uncover performance bottlenecks
  • 5+ years of professional software engineering experience, with at least 3 years in machine learning infrastructure such as developing, designing, scaling, training, deploying, and optimizing large-scale machine learning systems from data to model
  • Solid experience in the development and optimization of machine learning infrastructure tools like DeepSpeed, PyTorch, TensorFlow, Ray, or similar frameworks

This is an external listing. JobSpring does not represent or verify the employer. Report this listing