Skip to content
← Back to job listings

Machine Learning Engineer (GPU Kernel and Runtime)

Waymo · Mountain View, CA, United States

External listingfull-time8 days ago

About The Role

Join Waymo, a leader in autonomous driving technology, as a Machine Learning Engineer. In this role, you will work on the next generation of Waymo's onboard ML inference engine, collaborating with ML practitioners and analyzing ML workload performance at the hardware level. You will have the opportunity to optimize deep learning models for limited computation resources and develop tools for optimal resource usage and platform reliability. This position offers a hybrid work model, competitive compensation, and a comprehensive benefits package.

  • Collaborer avec des praticiens de l'apprentissage automatique sur des modèles pour la perception, la prédiction du comportement et la planification, afin de comprendre leurs modèles et de les accélérer à bord grâce au développement de noyaux GPU NVIDIA personnalisés.
  • Analyser les performances des charges de travail ML au niveau matériel ; appliquer des techniques manuelles et assistées par IA et développer des bibliothèques d'opérateurs CUDA/Triton personnalisées et hautement optimisées adaptées aux architectures spécifiques de Waymo.
  • Construire des outils pour évaluer, profiler l'exécution GPU et productiser des modèles d'apprentissage profond pour un déploiement à bord et hors bord rationalisé et robuste.
  • Strong C++ and CUDA programming skills
  • Extensive experience in NVIDIA GPU Kernel development to accelerate deep learning models
  • B.S. or M.S. in CS, EE, Deep Learning or a related field
  • Proven debugging and optimization experience on the XLA:GPU compiler, as well as the NVIDIA runtime stack
  • 5+ years of industry experience on system performance, hardware-level GPU optimization, or ML compilers
  • Passion for developing and optimizing ML software stacks for modern ML accelerator architectures (framework, runtime library, ML compiler, efficient deep learning etc.)
  • Custom GPU kernel development
  • ML runtime optimization
  • CUDA profiling & debugging
  • C++ Coding
  • Role-Related Knowledge
  • In-depth knowledge of ML frameworks, ML compilers, and IRs (Triton, HLO, MLIR, CuTe DSL, cuTile) or modern ML system architectures
  • Solid experience with designing, training and debugging deep learning models to achieve the highest scores/accuracies
  • Experience with advanced NVIDIA profiling (e.g., Nsight Compute) and debugging (e.g. cuda-gdb) tools
  • Strong Python programming skills
  • <M.Sc> or PhD in Computer Science, Mathematics or a related field

This is an external listing. JobSpring does not represent or verify the employer. Report this listing