← Back to job listings
TA
Machine Learning Engineer (Inference)
Together AI · San Francisco, United States
About The Role
Join Together AI as a Machine Learning Engineer focused on optimizing and enhancing the performance of our AI inference systems. Collaborate with AI researchers and engineers to create cutting-edge AI solutions, design and build production systems, and develop runtime inference services for large-scale AI applications. Enjoy competitive health insurance, dental and vision insurance, mental health support, income protection, retirement plans, and a flexible time off policy.
- Concevoir et construire des systèmes de production qui alimentent le moteur d'inférence de Together AI, en garantissant la fiabilité et la performance à grande échelle.
- Développer et optimiser les services d'inférence en temps d'exécution pour des applications d'IA à grande échelle.
- Collaborer avec des chercheurs, des ingénieurs, des chefs de produit et des designers pour apporter de nouvelles fonctionnalités et capacités de recherche au monde.
- If you are passionate about AI inference, PyTorch, and developing high-performance systems, we want to hear from you
- 3+ years of experience writing high-performance, well-tested, production-quality code
- Proficiency with Python and PyTorch
- Excellent understanding of low-level operating systems concepts including multi-threading, memory management, networking, storage, performance, and scale
- Preferred: Knowledge of existing AI inference systems such as TGI, vLLM, TensorRT-LLM, Optimum
- Preferred: Knowledge of CUDA/Triton programming
- Preferred: Knowledge of AI inference techniques such as speculative decoding
- Nice to have: Knowledge of Rust, Cython and compilers
- Demonstrated experience in building high performance libraries and tooling
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring