← Back to job listings
TA
Research Engineer (Core ML)
Together AI · San Francisco, United States
About The Role
Join Together AI as a Research Engineer (Core ML) and work at the intersection of efficient inference and post-training/RL systems. You will be responsible for designing and prototyping algorithms, implementing changes in high-performance inference engines, and optimizing performance across GPU, networking, and memory layers. This role requires strong coding ability in Python, experience in RL/post-training for large models, and a bias toward implementation and shipping.
- Conception et prototypage d'algorithmes, d'architectures et de stratégies de planification pour une inférence à faible latence et à haut débit.
- Mise en œuvre et maintenance des modifications dans les moteurs d'inférence haute performance, y compris les backends de noyau, le décodage spéculatif, la quantification, etc.
- Conception et exploitation de pipelines RL et post-formation, en optimisant conjointement les algorithmes et les systèmes.
- We don’t expect anyone to check every box below. People on this team typically have deep expertise in one or more areas and enough breadth (or interest) to work effectively across the stack
- The closer you are to full‑stack (inference + post‑training/RL + systems), the stronger the fit—but being spiky in one area and eager to grow is absolutely okay
- Strong coding ability in Python
- RL‑first profile: RL / post‑training for LLMs or large models (e.g., GRPO, RLHF/RLAIF, DPO‑like methods, reward modeling), and using these to train or fine‑tune real models
- Model architecture design for Transformers or other large neural nets
- Can read new RL / post‑training papers, understand their implications on the stack, and design minimal, correct changes in the right layer (training engine vs. inference engine vs. data / API)
- Able to take a new sampling method, scheduler, or RL update and turn it into a production‑grade implementation in the engine and/or training stack
- You enjoy collaborating with infra, research, and product teams, and you care about both scientific quality and user‑visible wins
- Have a bias toward implementation and shipping—you are excited to modify real engines and services, not just prototype in research code
- Systems‑first profile: Large‑scale inference systems (e.g., SGLang, vLLM, FasterTransformer, TensorRT, custom engines, or similar), GPU performance, distributed serving
- You naturally ask: “Where in the stack is this really bottlenecked?”
- Experience profiling and optimizing performance across GPU, networking, and memory layers
- Operate well as a full‑stack problem solver:
- Distributed systems / high‑performance computing for ML
- Have a solid research foundation in your area(s) of depth:
- Are comfortable working from algorithms to engines:
- Track record of impactful work in ML systems, RL, or large‑scale model training (papers, open‑source projects, or production systems)
- Have strong expertise in at least one of the following, and are excited to collaborate across (and grow into) the others:
- 3+ years of experience working on ML systems, large‑scale model training, inference, or adjacent areas (or equivalent experience via research / open source)
- Advanced degree in Computer Science, EE, or a related field, or equivalent practical experience
- Demonstrated experience owning complex technical projects end‑to‑end
- If you’re excited about the role and strong in some of these areas, we encourage you to apply even if you don’t meet every single requirement
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring