← Back to job listings
H1
Staff Data Engineer (Emerald)
H1 · New York, United States
About The Role
Join H1 as a Staff Data Engineer on the Emerald team, where you will shape the architecture and technical direction of our healthcare entity resolution platform. You will lead a small team of engineers while remaining hands-on technically, owning the systems and pipelines that power automatching, identity mapping, deduplication, and enrichment workflows. You will collaborate closely with various teams to improve platform accuracy, scalability, reliability, and operational efficiency.
- Lead the design, optimization, and scalability of distributed Spark/PySpark pipelines powering entity resolution and large-scale healthcare data processing.
- Own systems supporting automatching, identity mapping, grouping logic, deduplication, enrichment, and auto-approval workflows across healthcare provider and organization datasets.
- Drive infrastructure optimization initiatives focused on improving throughput, runtime, observability, and cloud compute cost efficiency.
- Experience with containerization and infrastructure technologies such as Docker, Kubernetes, and Terraform
- Strong grasp of software engineering fundamentals including distributed systems, data structures, concurrency, and system design
- Experience with orchestration and lakehouse technologies such as Argo and Hudi or comparable platforms
- Strong coding experience in Python (PySpark), Scala, Java, or equivalent languages used for distributed processing systems
- Experience working with relational or distributed databases such as PostgreSQL or Redshift
- Deep expertise with distributed data processing frameworks such as Apache Spark and Hadoop, particularly within AWS environments
- Experience performing root cause analysis across large-scale distributed systems and complex data pipelines
- Demonstrated technical leadership experience mentoring engineers and driving complex technical initiatives
- Extensive experience with Apache Spark and AWS-based big data technologies including EMR, S3, and distributed compute environments
- Experience with streaming and event-driven architectures using technologies such as Kafka or Spark Streaming
- 8+ years of experience building and maintaining large-scale distributed data systems and pipelines
- Proven ability to operate effectively within highly scalable, production-grade distributed systems
- Strong proficiency in Python (PySpark), Scala, Java, or other modern programming languages used for large-scale distributed processing
- Experience with entity resolution, identity mapping, automatching, deduplication, or large-scale matching systems is strongly preferred
- You bring strong hands-on engineering expertise across distributed computing, large-scale data processing, and infrastructure optimization while also helping guide technical direction and mentor engineers across the organization
- Experience optimizing large-scale Spark workloads for performance, scalability, and infrastructure cost efficiency
- Ability to write clean, maintainable, modular, and production-grade code
- Experience building scalable ETL/ELT frameworks across both batch and streaming architectures
- Strong understanding of distributed file formats including Apache Parquet and Apache AVRO
- Experience working with healthcare, life sciences, Real World Evidence (RWE), or large-scale healthcare datasets is strongly preferred
- Experience improving performance, scalability, observability, and infrastructure efficiency within distributed systems
- Strong communication and collaboration skills across both technical and non-technical stakeholders
- Familiarity with modern development and infrastructure tooling including Git, CI/CD pipelines, Docker, Kubernetes, Terraform, Argo, Hudi, and JIRA
- Experience with streaming technologies such as Kafka, Spark Streaming, or KSQL
- You are an experienced data engineer with deep expertise building and optimizing distributed data systems in cloud-native environments. You thrive solving complex scalability and performance challenges across high-volume data processing systems and enjoy operating in highly technical, fast-paced engineering environments
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring