Skip to content
← Back to job listings

AI Inference Internship

Perplexity AI · London, United Kingdom

External listinginternshipabout 1 month ago

About The Role

Join our AI Inference team as an intern and work on optimizing the performance of our models and inference engine. You will have the opportunity to improve serving latency and throughput, support new models, and optimize inference across the entire stack. This is a 13-week internship, full-time or part-time, in-person in our London office. Candidates should be pursuing a Master's or PhD in Computer Science with a focus on performance-related subjects and have experience with ML frameworks, high-performance computing, and GPU programming.

  • Collaborate with the inference team to enhance serving latency and throughput for AI models.
  • Assist in bringing up support for new models and implementing state-of-the-art inference optimizations.
  • Optimize inference across the entire stack, from GPU kernels to serving endpoints, ensuring high performance.
  • Pursuing a Master's or PhD in Computer Science with a focus on performance-related subjects (HPC, Compilers, Distributed Systems)
  • Experience with ML frameworks (Torch, JAX)
  • Experience with High-Performance Computing (OpenMPI)
  • Experience with GPU programming (CUDA, Triton)
  • Strong engineering track record with proven knowledge of fundamentals and programming languages (multi-threaded programming, networking, compilation, systems programming, etc)

This is an external listing. JobSpring does not represent or verify the employer. Report this listing