Skip to content
← Back to job listings

AI Researcher (Core Machine Learning, Turbo)

Together AI · San Francisco, United States

External listingfull-time23 days ago

About The Role

Join the Turbo team at Together, where you'll work at the intersection of efficient inference and post-training/RL systems. Your role will involve designing and prototyping algorithms, implementing changes in high-performance inference engines, optimizing performance, and operating RL and post-training pipelines. You'll collaborate across the stack, from RL algorithms and training engines to kernels and serving systems, and drive roadmap items that require real engine modification. This position offers competitive health insurance, flexible time off, and a supportive work environment.

  • Design and prototype algorithms, architectures, and scheduling strategies for low-latency, high-throughput inference.
  • Implement and maintain changes in high-performance inference engines, including kernel backends, speculative decoding, quantization, etc.
  • Profile and optimize performance across GPU, networking, and memory layers to improve latency, throughput, and cost.
  • Track record of impactful work in ML systems, RL, or large‑scale model training (papers, open‑source projects, or production systems)
  • Experience profiling and optimizing performance across GPU, networking, and memory layers
  • Are comfortable working from algorithms to engines:
  • You enjoy collaborating with infra, research, and product teams, and you care about both scientific quality and user‑visible wins
  • Systems‑first profile: Large‑scale inference systems (e.g., SGLang, vLLM, FasterTransformer, TensorRT, custom engines, or similar), GPU performance, distributed serving
  • Model architecture design for Transformers or other large neural nets
  • Have strong expertise in at least one of the following, and are excited to collaborate across (and grow into) the others:
  • Able to take a new sampling method, scheduler, or RL update and turn it into a production‑grade implementation in the engine and/or training stack
  • Have a solid research foundation in your area(s) of depth:
  • RL‑first profile: RL / post‑training for LLMs or large models (e.g., GRPO, RLHF/RLAIF, DPO‑like methods, reward modeling), and using these to train or fine‑tune real models
  • Distributed systems / high‑performance computing for ML
  • Operate well as a full‑stack problem solver:
  • Strong coding ability in Python
  • Can read new RL / post‑training papers, understand their implications on the stack, and design minimal, correct changes in the right layer (training engine vs. inference engine vs. data / API)
  • You naturally ask: “Where in the stack is this really bottlenecked?”
  • Advanced degree in Computer Science, EE, or a related field, or equivalent practical experience
  • If you’re excited about the role and strong in some of these areas, we encourage you to apply even if you don’t meet every single requirement
  • 3+ years of experience working on ML systems, large‑scale model training, inference, or adjacent areas (or equivalent experience via research / open source)
  • Demonstrated experience owning complex technical projects end‑to‑end

This is an external listing. JobSpring does not represent or verify the employer. Report this listing