Skip to content
← Back to job listings

Staff Research Engineer (Post-training & Evaluation)

Reddit · United States

External listingfull-timeabout 1 month ago

About The Role

Join Reddit's AI Engineering team as a Staff Research Engineer focused on Post-Training & Evaluation Science. You'll be responsible for defining the "Reddit Benchmark" evaluation standard, ensuring model quality across various metrics, and driving the practice of evaluation as a release gate. You'll also design post-training recipes, evaluate checkpoints, and partner with Safety Engineering. This role requires a PhD or MS in a related field, fluency in Python, and 6+ years of professional ML experience.

  • Définir la norme d'évaluation "Reddit Benchmark" : Posséder la méthodologie pour mesurer rigoureusement la qualité des modèles.
  • Assurer la fiabilité de l'évaluation et la rigueur statistique : Établir la science derrière des évaluations fiables.
  • Concevoir des recettes et des stratégies de post-formation : Concevoir des recettes SFT qui convertissent les modèles de base en points de terminaison utiles.
  • PhD or MS in CS, ML, NLP, IR, or a related quantitative field — or equivalent industry research experience
  • Fluency in Python; strong data-pipeline and eval-harness engineering (e.g., Hugging Face Transformers, vLLM, lm-eval-harness). Working knowledge of PyTorch and distributed training (FSDP2, DeepSpeed ZeRO-3) sufficient to direct and debug post-training runs
  • Experience evaluating both generation and representation/classification: model-as-a-judge for generative quality and precision/recall, PR-AUC, retrieval/MTEB-style metrics, gold-label denoising, and label-noise handling
  • Strong experience building custom, domain-specific evaluation harnesses (e.g., lm-eval-harness, Inspect AI, LightEval) — you know the strengths and limits of benchmarks like MMLU and GSM8K and when they don't apply, and you treat eval sets as versioned, frozen, regression-tracked code
  • Deep expertise in evaluation reliability: judge/sample variance, multi-sample scoring, calibration, statistical significance, and the failure modes of automated evaluation
  • Deep understanding of Continuous Pre-training (CPT), Instruction Tuning (SFT), and how data quality shapes model behavior
  • 6+ years of professional ML experience (or PhD + 4+) with a direct focus on LLM post-training and evaluation
  • Experience with MLflow or similar experiment-tracking frameworks
  • Familiarity with modern fine-tuning frameworks (Axolotl, TorchTune) and PyTorch-native training stacks (TorchTitan)
  • Synthetic data generation techniques (e.g., Self-Instruct)
  • Experience with preference optimization (DPO, RLHF, RLAIF, GRPO)
  • Publications in NLP/ML/FAccT or related venues, or other evidence of research leadership
  • Experience evaluating multimodal models (embeddings, hateful-memes-style classification)

This is an external listing. JobSpring does not represent or verify the employer. Report this listing