← Back to job listings
NE
Machine Learning Engineer (LLM Post-Training)
NewsBreak · Mountain View, CA, United States
About The Role
Join our team as a Machine Learning Engineer focused on the post-training of large language models. In this hands-on role, you will own the full post-training stack, including continuous pre-training, supervised fine-tuning, and reinforcement learning. You will work closely with product and business teams to translate real-world use cases into concrete training objectives and deliver model improvements quickly. This is a high-ownership role for someone with practical experience in training models.
- Lead the post-training of large language models (LLMs) across the full pipeline, with a primary focus on reinforcement learning (RL).
- Design, build, and curate the data that drives each training stage, and define data-preparation strategies tailored to specific business needs.
- Partner closely with business and product stakeholders to understand their scenarios, rapidly convert requirements into training plans, and deliver targeted model capabilities.
- Strong PyTorch fundamentals; working familiarity with frameworks such as Hugging Face TRL/Accelerate, DeepSpeed or FSDP, and inference engines like vLLM
- Strong data engineering for ML. You can independently design data-preparation plans for a given business scenario — sourcing, cleaning, filtering, labeling strategy, and synthetic/preference data generation — to meet specific product requirements
- A bias toward fast iteration and business impact, with strong communication skills to work across research and product teams
- Proven large-scale GPU training ability. You have trained LLMs on mid-to-large GPU hardware and are comfortable with distributed training and debugging at scale
- Hands-on LLM post-training experience. You have personally run CPT, SFT, and RL training — with demonstrated, practical RL experience (RLHF / PPO / GRPO / DPO or similar), beyond just launching training scripts
- Solid understanding of tokenization, attention, chat templates, and common failure modes in alignment/agent training
- Experience designing reward models or rule-based verifiers for RL
- Publications or open-source contributions in LLM post-training or RL
- Experience with tool-use / agentic model training (function calling, multi-step planning)
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring