Skip to content
← Back to job listings

Machine Learning Engineer (Inference)

Together AI · San Francisco, United States

External listingfull-time22 days ago

About The Role

Join Together AI as a Machine Learning Engineer focused on optimizing and enhancing the performance of our AI inference systems. Collaborate with AI researchers and engineers to create cutting-edge AI solutions, design and build production systems, and develop runtime inference services for large-scale AI applications. Enjoy competitive health insurance, dental and vision insurance, mental health support, income protection, retirement plans, and a flexible time off policy.

  • Concevoir et construire des systèmes de production qui alimentent le moteur d'inférence de Together AI, en garantissant la fiabilité et la performance à grande échelle.
  • Développer et optimiser les services d'inférence en temps d'exécution pour des applications d'IA à grande échelle.
  • Collaborer avec des chercheurs, des ingénieurs, des chefs de produit et des designers pour apporter de nouvelles fonctionnalités et capacités de recherche au monde.
  • If you are passionate about AI inference, PyTorch, and developing high-performance systems, we want to hear from you
  • 3+ years of experience writing high-performance, well-tested, production-quality code
  • Proficiency with Python and PyTorch
  • Excellent understanding of low-level operating systems concepts including multi-threading, memory management, networking, storage, performance, and scale
  • Preferred: Knowledge of existing AI inference systems such as TGI, vLLM, TensorRT-LLM, Optimum
  • Preferred: Knowledge of CUDA/Triton programming
  • Preferred: Knowledge of AI inference techniques such as speculative decoding
  • Nice to have: Knowledge of Rust, Cython and compilers
  • Demonstrated experience in building high performance libraries and tooling

This is an external listing. JobSpring does not represent or verify the employer. Report this listing