← Back to job listings
H1
Staff Data Engineer (Data Lake)
H1 · New York, United States
About The Role
Join H1 as a Staff Data Engineer on the Data Lake team, where you will shape the architecture and direction of our core data platform. You will architect and scale distributed ETL/ELT pipelines, improve data quality workflows, and optimize data processing frameworks. This role offers the opportunity for technical leadership and mentorship, as well as the chance to influence technical direction and grow into broader engineering leadership responsibilities.
- Architect, build, and scale distributed ETL/ELT pipelines and large-scale ingestion frameworks across structured and unstructured healthcare datasets.
- Lead the evolution of H1’s Data Lake architecture with a focus on scalability, observability, reliability, and cost optimization.
- Own and improve data quality, validation, normalization, and standardization workflows across thousands of global data sources.
- You are a highly technical data engineer who thrives in lean, high-ownership environments and enjoys solving complex distributed systems challenges. You are excited by the opportunity to influence technical direction, mentor engineers, and grow into broader engineering leadership responsibilities while remaining hands-on
- Experience with large-scale data cleaning, parsing, normalization, and validation workflows preferred
- Experience with orchestration and workflow management tools such as Argo, Airflow, or similar technologies
- You are comfortable leading technical initiatives and influencing architecture decisions across teams
- Demonstrated technical leadership experience with interest in or experience mentoring and leading engineers
- Deep experience with Apache Spark and cloud-native big data platforms, preferably within AWS environments (EMR, Glue, S3, Athena, Redshift, or similar)
- Experience implementing monitoring, observability, and data quality frameworks within production environments
- Experience using AI-assisted coding tools (e.g., GitHub Copilot, Claude Code) to accelerate development while maintaining quality is encouraged
- Experience designing and scaling modern cloud-native data lake architectures and large-scale ingestion frameworks
- Exposure to ML/AI-driven data enrichment, parsing, or validation workflows is a plus
- You communicate effectively with both technical and non-technical stakeholders
- Advanced SQL expertise, including performance tuning and optimization across large datasets
- You are energized by ownership, autonomy, and solving ambiguous technical challenges
- 8+ years of experience in data engineering, software engineering, or related fields with significant experience building and scaling distributed data platforms
- Strong understanding of distributed storage systems, partitioning strategies, and file formats such as Parquet, Avro, and ORC
- Experience with Docker, Kubernetes, and modern containerization technologies
- You have deep experience designing and scaling distributed data platforms and large-scale pipelines in cloud-native environments
- You excel at building reliable, observable, and maintainable data systems supporting critical business and analytics workloads
- Experience working with healthcare, life sciences, publication, or large-scale entity-resolution datasets preferred
- You have strong expertise in distributed processing, performance optimization, and modern data architecture patterns
- Strong proficiency in Python (PySpark), Java, Scala, or similar programming languages
- You enjoy mentoring engineers and helping raise the engineering bar across teams
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring