Skip to content
← Back to job listings

Staff Machine Learning Performance Engineer (Inference Optimisation)

Wayve · London, United Kingdom

External listingfull-timeabout 1 month ago

About The Role

Join Wayve, a pioneering company in the field of self-driving technology. As a Staff Machine Learning Performance Engineer, you will be instrumental in optimizing ML inference for edge accelerators and GPUs. Your work will directly impact the efficiency of large transformer-based models on low-cost, low-power edge devices, enabling Wayve's first driving product. This hands-on role involves profiling and pinpointing bottlenecks across the full inference stack, implementing and validating optimizations, and collaborating with model developers to influence architecture and deployment decisions. You will also contribute to technical roadmaps and tooling, raising the standard of performance engineering across the team.

  • Profiling and pinpointing bottlenecks across the full inference stack (model graph, compiler/runtime, kernel execution, memory movement) and delivering measurable improvements.
  • Implementing and validating optimisations in compilers, runtimes, and/or kernels (e.g. operator fusion, scheduling, quantisation-aware performance, custom kernels).
  • Building robust benchmarking and regression testing to ensure performance improvements hold across models, devices, and software releases.
  • Proven experience improving performance in production systems with tight constraints (latency, memory, bandwidth, power/thermal, or cost)
  • Strong software engineering fundamentals (debugging, profiling, testing, and maintainable code)
  • Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL) and confidence learning adjacent frameworks quickly
  • Comfort operating at multiple levels of abstraction — from high-level model behaviour down to low-level kernel/runtime execution
  • Clear communicator and collaborative teammate; able to align multiple stakeholders on performance trade-offs and priorities
  • Exposure to embedded or edge deployment of ML models, including benchmarking on real devices and handling system-level constraints
  • Experience with NVIDIA and/or Qualcomm SoCs and performance tooling
  • Experience mentoring others and/or driving technical direction in a small, fast-moving team
  • Python and C++ proficiency
  • We understand that everyone has a unique set of skills and experiences and that not everyone will meet all of the requirements listed above. If you’re passionate about self-driving cars and think you have what it takes to make a positive impact on the world, we encourage you to apply

This is an external listing. JobSpring does not represent or verify the employer. Report this listing