Skip to content
← Back to job listings

Machine Learning Engineer (LLM Post-Training)

NewsBreak · Mountain View, CA, United States

External listingfull-timeabout 2 months ago

About The Role

Join our team as a Machine Learning Engineer focused on the post-training of large language models. In this hands-on role, you will own the full post-training stack, including continuous pre-training, supervised fine-tuning, and reinforcement learning. You will work closely with product and business teams to translate real-world use cases into concrete training objectives and deliver model improvements quickly. This is a high-ownership role for someone with practical experience in training models.

  • Lead the post-training of large language models (LLMs) across the full pipeline, with a primary focus on reinforcement learning (RL).
  • Design, build, and curate the data that drives each training stage, and define data-preparation strategies tailored to specific business needs.
  • Partner closely with business and product stakeholders to understand their scenarios, rapidly convert requirements into training plans, and deliver targeted model capabilities.
  • Strong PyTorch fundamentals; working familiarity with frameworks such as Hugging Face TRL/Accelerate, DeepSpeed or FSDP, and inference engines like vLLM
  • Strong data engineering for ML. You can independently design data-preparation plans for a given business scenario — sourcing, cleaning, filtering, labeling strategy, and synthetic/preference data generation — to meet specific product requirements
  • A bias toward fast iteration and business impact, with strong communication skills to work across research and product teams
  • Proven large-scale GPU training ability. You have trained LLMs on mid-to-large GPU hardware and are comfortable with distributed training and debugging at scale
  • Hands-on LLM post-training experience. You have personally run CPT, SFT, and RL training — with demonstrated, practical RL experience (RLHF / PPO / GRPO / DPO or similar), beyond just launching training scripts
  • Solid understanding of tokenization, attention, chat templates, and common failure modes in alignment/agent training
  • Experience designing reward models or rule-based verifiers for RL
  • Publications or open-source contributions in LLM post-training or RL
  • Experience with tool-use / agentic model training (function calling, multi-step planning)

This is an external listing. JobSpring does not represent or verify the employer. Report this listing