Skip to content
← Back to job listings

Senior Machine Learning Engineer (Runtime and Serving)

Waymo · Mountain View, CA, United States

External listingfull-time22 days ago

About The Role

Join Waymo, a leader in autonomous vehicle technology, as a Senior Machine Learning Engineer. In this role, you will architect and develop a high-performance ML runtime and serving system, lead integration and feature development for ML inference runtimes, and drive the strategic migration of ML workloads. You will collaborate with world-class ML practitioners and design robust tooling for profiling and benchmarking. This position offers a hybrid work model, competitive compensation, and a comprehensive benefits package.

  • Architect and develop an efficient, high-performance ML runtime and serving system tailored for both onboard autonomous vehicle compute and large-scale, offboard data center environments.
  • Lead the integration and feature development for ML inference runtimes across both domains, balancing the strict real-time latency and memory constraints of onboard systems with the high-throughput, highly concurrent demands of offboard serving fleets.
  • Drive the strategic migration of ML workloads toward a JAX-native runtime architecture, which includes extending and modifying underlying ML compilers and runtimes (e.g., OpenXLA/PjRT, TensorRT).
  • PhD in CS, EE, Deep Learning or a related field
  • Experience modifying ML compilers, runtimes, or inference engines (e.g., TensorRT, ONNX Runtime, OpenXLA/PjRT, TVM)
  • Experience optimizing ML software for hardware accelerators (e.g., GPUs, TPUs, custom silicon)
  • Experience building low-latency, highly concurrent distributed backend systems
  • B.S. or M.S. in CS, EE, Deep Learning or a related field
  • 3+ years of production experience in Python and major deep learning frameworks (e.g., PyTorch, JAX)
  • 5+ years production programming in C++
  • 5+ years of professional software engineering experience focused on building, scaling, or maintaining ML systems and infrastructure
  • Experience building or scaling LLM serving systems, including expertise in distributed inference and performance optimization (e.g., KV/prefix caching, continuous batching)
  • Experience with custom kernel development (e.g., CUDA/CUDA Tile, Triton, JAX/Pallas)
  • Experience architecting unified serving APIs and optimizing tensor buffer management (e.g., zero-copy data transfer, shared memory) for complex, multi-model inference pipelines

This is an external listing. JobSpring does not represent or verify the employer. Report this listing