← Back to job listings
RE
Staff Research Engineer (Post-training & Evaluation)
Reddit · United States
About The Role
Join Reddit's AI Engineering team as a Staff Research Engineer focused on Post-Training & Evaluation Science. You'll be responsible for defining the "Reddit Benchmark" evaluation standard, ensuring model quality across various metrics, and driving the practice of evaluation as a release gate. You'll also design post-training recipes, evaluate checkpoints, and partner with Safety Engineering. This role requires a PhD or MS in a related field, fluency in Python, and 6+ years of professional ML experience.
- Définir la norme d'évaluation "Reddit Benchmark" : Posséder la méthodologie pour mesurer rigoureusement la qualité des modèles.
- Assurer la fiabilité de l'évaluation et la rigueur statistique : Établir la science derrière des évaluations fiables.
- Concevoir des recettes et des stratégies de post-formation : Concevoir des recettes SFT qui convertissent les modèles de base en points de terminaison utiles.
- PhD or MS in CS, ML, NLP, IR, or a related quantitative field — or equivalent industry research experience
- Fluency in Python; strong data-pipeline and eval-harness engineering (e.g., Hugging Face Transformers, vLLM, lm-eval-harness). Working knowledge of PyTorch and distributed training (FSDP2, DeepSpeed ZeRO-3) sufficient to direct and debug post-training runs
- Experience evaluating both generation and representation/classification: model-as-a-judge for generative quality and precision/recall, PR-AUC, retrieval/MTEB-style metrics, gold-label denoising, and label-noise handling
- Strong experience building custom, domain-specific evaluation harnesses (e.g., lm-eval-harness, Inspect AI, LightEval) — you know the strengths and limits of benchmarks like MMLU and GSM8K and when they don't apply, and you treat eval sets as versioned, frozen, regression-tracked code
- Deep expertise in evaluation reliability: judge/sample variance, multi-sample scoring, calibration, statistical significance, and the failure modes of automated evaluation
- Deep understanding of Continuous Pre-training (CPT), Instruction Tuning (SFT), and how data quality shapes model behavior
- 6+ years of professional ML experience (or PhD + 4+) with a direct focus on LLM post-training and evaluation
- Experience with MLflow or similar experiment-tracking frameworks
- Familiarity with modern fine-tuning frameworks (Axolotl, TorchTune) and PyTorch-native training stacks (TorchTitan)
- Synthetic data generation techniques (e.g., Self-Instruct)
- Experience with preference optimization (DPO, RLHF, RLAIF, GRPO)
- Publications in NLP/ML/FAccT or related venues, or other evidence of research leadership
- Experience evaluating multimodal models (embeddings, hateful-memes-style classification)
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring