← Back to job listings
HT
Senior Data Engineer
Hack The Box · United States
About The Role
Join Hack The Box as a Senior Data Engineer, where you'll own and evolve our data pipelines on GCP. You'll work end-to-end on streaming and batch pipelines, design ELT/ETL processes, and partner with ML engineers. This role offers a chance to drive the migration off Snowflake onto our GCP-native stack and continuously improve data quality, reliability, and cost efficiency. Enjoy benefits like 25 annual leave days, private insurance, and a dedicated budget for training and professional development.
- Posséder et faire évoluer les pipelines de données sur GCP, en construisant de nouveaux pipelines, en renforçant les existants et en améliorant la qualité des données.
- Travailler de bout en bout sur les pipelines de streaming et de batch, de l'ingestion d'événements à la transformation, au service et à la couche de fonctionnalités qui alimente nos produits ML et AI.
- Collaborer avec les ingénieurs ML sur les pipelines de fonctionnalités, surveiller la dérive des données et maintenir les modèles bien alimentés et réentraînés.
- Solid SQL and strong Python — you write production-quality code, not just notebooks
- Working knowledge of ML in production — feature engineering, feature stores, model deployment, drift monitoring, retraining
- Workflow orchestration experience with Airflow (or Prefect/Dagster)
- Experience with ClickHouse or another columnar OLAP engine in production
- CI/CD mindset, infrastructure-as-code sensibility, and a bias for simple, observable systems
- Bonus: CDC tooling (Datastream, Debezium), Vertex AI / Feature Store
- Docker & Kubernetes experience
- Strong data modelling and warehouse architecture skills (dimensional modelling, event-driven, lakehouse patterns)
- Comfortable with dbt or equivalent transformation frameworks
- Production experience with streaming pipelines on Dataflow/Beam, Flink, or Spark Structured Streaming, ingesting from Kafka and/or Pub/Sub
- Experience migrating off legacy warehouses (Snowflake, Redshift, Synapse) onto cloud-native stacks is a plus
- Hands-on experience with GCP data services — BigQuery is a must; Pub/Sub, Dataflow, Bigtable, Cloud Composer are strong pluses
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring