Skip to content
← Back to job listings

Staff Applied Scientist (Reinforcement Learning)

Hippocratic AI · Menlo Park, CA, United States

External listingfull-time5 days ago

About The Role

Join our team as a Staff Applied Scientist specializing in Reinforcement Learning (RL) and On-Policy Distillation (OPD) post-training. You will be responsible for improving our models' clinical reasoning, safety, and alignment, with the potential to impact millions of patients across diverse clinical use cases. Your work will involve designing RL and OPD post-training methods, building and evaluating reward models, developing conversational AI environments for healthcare RL training, automating post-training loops, and running rigorous experiments. You will collaborate with research, engineering, and clinical teams.

  • Ownership of the Reinforcement Learning (RL) and On-Policy Distillation (OPD) post-training pipeline to enhance clinical reasoning and safety.
  • Designing and implementing RL and OPD post-training methods, including RLHF, RLVR, and OPD.
  • Collaborating with research, engineering, and clinical teams to automate post-training loops and run rigorous experiments.
  • Experience with RLHF, RLVR, LLM-as-judge or similar methods for LLM post-training
  • MS or PhD in CS or relevant field
  • 5+ years or experience in NLP, LLM training, or RL
  • 2+ years experience in RL for LLM post-training
  • Strong Python and PyTorch coding skills
  • Experience with large-scale (50B+ parameter and multi-node) LLM training
  • Publications at top venues (NeurIPS, ICML, ICLR, ACL, EMNLP)
  • Healthcare domain experience

This is an external listing. JobSpring does not represent or verify the employer. Report this listing