Skip to content
← Back to job listings

Research Engineer (Domain Scaling)

Normal Computing · New York, United States

External listingfull-time14 days ago

About The Role

Join our team as a Founding Data Engineer, where you will play a crucial role in accelerating the design and verification of silicon. You will be responsible for generating synthetic training data, mining our agents' runs for high-quality trajectories, and negotiating access to real customer data. You will work closely with verification engineers, ML engineers, and pipeline engineers, and will have ownership of the data strategy. Your main goal will be to improve our models' performance in hardware design, verification, and EDA workflows.

  • Ownership of the end-to-end pipeline for generating synthetic training data, mining agent runs, and negotiating access to real customer data.
  • Collaboration with verification engineers, ML engineers, and pipeline engineers to improve models and ensure high-quality data production.
  • Management of the data acquisition strategy, including identifying, evaluating, and acquiring datasets relevant to hardware design and verification.
  • You approach data acquisition as an engineering problem: systematic, measurable, and outcome-driven
  • You're organized and documentation-minded: you track provenance, ownership, and lineage as a matter of habit
  • You can evaluate data quality independently, spotting noise, bias, and gaps without needing someone to tell you what to look for
  • You've built or used a data flywheel: model outputs, curated, into the next training round
  • You're comfortable working across multiple technical roles and synthesizing feedback from domain experts, ML engineers, and pipeline engineers
  • You've shipped a synthetic-data or training-data pipeline that produced a measurable downstream model improvement you can describe by number, not vibes
  • Experience acquiring data from a variety of sources, both paid and unpaid, and managing vendor relationships
  • Familiarity with SystemVerilog, Verilog, and UVM
  • Background in code-model or agent training-data pipelines (e.g. SWE-bench-style data, code-model post-training)
  • Prior work in a startup or fast-moving research environment where the data strategy was still being defined
  • Experience with automated data collection, web scraping, or corpus curation at scale

This is an external listing. JobSpring does not represent or verify the employer. Report this listing