Skip to content
← Back to job listings

Research Engineer

Super Annotate · San Francisco, United States

External listingfull-time6 days ago

About The Role

Join our expanding research team as a Research Engineer. You will take a research direction and independently identify supporting resources, implement relevant methods, and build a process to reproduce and improve prior work. You will own projects end to end, partner with strategic project leads and technical leads, validate ideas through hands-on implementation, and turn research directions into tangible outputs. You should have hands-on experience with RL/agentic systems, AI/ML evaluation and benchmarking, or multimodal ML, strong Python skills, and a MS or PhD in ML, CS, or a related quantitative field.

  • Research and implement relevant methods and benchmarks to support research directions.
  • Own projects end to end, including scoping, MVP implementation, and validation.
  • Translate ambiguous requirements into a concrete, testable research plan in collaboration with project leads.
  • Hands-on experience with at least one of: RL/agentic systems, AI/ML evaluation and benchmarking, or multimodal ML
  • Strong Python and the engineering ability to build and ship your own experiments – eval harnesses, environments, infrastructure – without relying on a platform team
  • Real ML depth: you understand how models are trained and evaluated, not just how to call an API. You can read a paper, judge whether its claims hold, and reimplement the method
  • MS or PhD in ML, CS, or a related quantitative field – or equivalent demonstrated research experience (publications, significant open-source research work, industry research)
  • High autonomy: you can turn an ambiguous direction into a concrete research plan and notice when something's off before being told
  • Clear technical writing
  • Publication track record (first-author preferred)
  • Experience with agent or multimodal benchmarks (OSWorld, MMMU, WebArena, SWE-bench, or similar) or building RL environments/gyms
  • Familiarity with reward modeling, reward hacking, or verifier/judge reliability
  • Familiarity with synthetic data generation or human-in-the-loop (HITL) workflows
  • A deep RL background specifically
  • Experience with cloud infrastructure and containerized environments

This is an external listing. JobSpring does not represent or verify the employer. Report this listing