← Back to job listings
WA
Technical Lead Manager (Machine Learning Runtime & Serving)
Waymo · Mountain View, CA, United States
About The Role
Waymo is seeking a senior Technical Lead Manager (TLM) Machine Learning Engineer to guide the technical vision of their core ML infrastructure. In this role, you will manage a team of 6 engineers and architect scalable, high-performance ML runtime systems. You will navigate complex engineering trade-offs, spearhead the transition to a JAX-native runtime architecture, and drive systemic performance excellence. The position offers a range of benefits, including medical, dental, and vision insurance, competitive compensation, and a hybrid work model.
- Guiding the technical vision of Waymo's core ML infrastructure and managing a high-performing team of engineers.
- Architecting scalable, high-performance ML runtime systems that operate across constrained edge compute environments and large-scale data centers.
- Driving the strategic transition of core ML workloads to a JAX-native runtime architecture and partnering with ML researchers to optimize performance.
- B.S. or M.S. in CS, EE, Deep Learning or a related field
- People management experience, with a proven track record of recruiting, mentoring, and guiding high-performing teams of senior engineers
- Proven track record of optimizing ML software to maximize the performance of hardware accelerators (e.g., GPUs, TPUs, or custom silicon)
- Strong production programming expertise
- 8+ years of professional software engineering experience architecting, building, and scaling complex ML systems and infrastructure
- Hands-on experience developing distributed backend systems that are low-latency, highly concurrent, and fault-tolerant at scale
- PhD in CS, EE, Deep Learning or a related field
- Deep expertise in modifying and extending ML software stacks, including compilers, runtimes, or inference engines (e.g., OpenXLA/PjRT, TensorRT, ONNX Runtime, TVM)
- Strong background in building and scaling LLM serving systems, leveraging advanced distributed inference and performance optimization techniques
- Deep expertise in edge computing and automotive ML deployment, navigating strict power, thermal, and real-time latency constraints to optimize and deploy mission-critical models on resource-constrained embedded hardware
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring