Skip to content
← Back to job listings

Data Engineer

Reality Defender · United States

External listingfull-time10 days ago

About The Role

Join Reality Defender as a Data Engineer, where you'll build and scale the infrastructure for our data platform. You'll design and operate pipelines for multi-terabyte and streaming datasets, work closely with ML engineers and researchers, and deploy and troubleshoot containerized workloads on Kubernetes and AWS. Enjoy benefits such as healthcare coverage, dental and vision plans, disability and life insurance, a learning and development budget, and 20 days of PTO per year.

  • Design, build, and operate large-scale data processing pipelines handling multi-terabyte and streaming datasets, including audio/video transcoding, feature extraction, and preprocessing workflows.
  • Deploy, scale, and troubleshoot containerized workloads on Kubernetes and AWS in production environments, ensuring reliability and performance.
  • Partner with ML engineers and researchers to support training pipelines, model retraining triggers, feature stores, and other MLOps workflows.
  • Strong programming skills in Python and SQL; experience with Golang is a plus
  • Proficiency with high-performance/distributed computing frameworks such as Spark and Ray for processing large-scale datasets
  • Experience designing and maintaining job orchestration systems, including dependency management, retries, monitoring, and alerting for production pipelines
  • Demonstrated track record building and operating large-scale data processing pipelines, ideally handling multi-terabyte or streaming datasets
  • Experience working with audio or video data at scale is a strong plus (e.g., transcoding, feature extraction, or preprocessing pipelines)
  • Hands-on experience with Kubernetes and AWS, including deploying, scaling, and troubleshooting containerized workloads in production environments
  • Familiarity with common data transformation patterns applied to large datasets (ETL/ELT, batch and stream processing, data validation and quality checks)
  • Bonus: experience orchestrating machine learning workflows (training pipelines, model retraining triggers, feature stores, or MLOps tooling)
  • Experience with workflow orchestration tools such as Airflow (or comparable systems like Dagster, Prefect, or Luigi) to schedule and manage complex data pipelines

This is an external listing. JobSpring does not represent or verify the employer. Report this listing