Staff Machine Learning Software Engineer (Infrastructure, Driver Understanding and Evaluation)
Waymo · Mountain View, CA, United States
About The Role
Join Waymo, a leader in autonomous vehicle technology. As a Staff Machine Learning Software Engineer, you will provide technical leadership on large-scale ML model architectures, build scalable systems for training and fine-tuning models, and oversee the production and optimization of machine learning models. You will collaborate cross-functionally to derive performance and system-level requirements for large ML systems and translate product/business goals into measurable technical deliverables. Enjoy a comprehensive benefits package, including medical, dental, and vision insurance, competitive compensation, and a hybrid work model.
- Provide deep technical leadership on large-scale ML model architectures, especially for autonomous vehicle models.
- Build scalable systems for training and fine-tuning large-scale models to evaluate interesting driving behaviors.
- Oversee the production and optimization of machine learning models aiming to assess Waymo’s expansive fleet of vehicles.
- 7+ years of professional software engineering experience, with at least 3 years in machine learning infrastructure such as developing, designing, scaling, training, deploying, and optimizing large-scale machine learning systems from data to model
- Deep understanding of state-of-the-art machine learning models such as autoregressive transformers
- Strong leadership skills with experience navigating cross-functional teams and providing technical leadership projects across multiple organizations
- A history of contributions to machine learning tooling and frameworks e.g. PyTorch, Jax, Tensorflow, Ray, or similar. The candidate should understand both the user facing API and the internal workings
- M.S. or Ph.D. degree Computer Science, Machine Learning, Artificial Intelligence, or a related technical field, or equivalent practical experience
- Strong expertise in distributed training techniques, including gradient sharding and optimization strategies for scaling large models across ML accelerator profiling tools to uncover performance bottlenecks
- 10+ years of professional software engineering experience, with at least 5 years in machine learning infrastructure such as developing, designing, scaling, training, deploying, and optimizing large-scale machine learning systems from data to model
- Deep understanding of state-of-the-art RL techniques, including those used for fine-tuning large models (e.g., from human feedback/preferences)
- Familiarity with large-scale simulation platforms and their integration with ML training workflows
- Experience in the autonomous vehicles domain, robotics, or complex simulation environments
- Experience designing and using metrics for evaluating complex AI systems
- Track record of technical leadership, influencing senior stakeholders, and driving innovation across team boundaries
- Excellent communication skills, with the ability to articulate complex technical concepts clearly
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring