Staff Backend Engineer (Catalog)
DataHub · Palo Alto, United States
About The Role
Join our team as a Staff Backend Engineer (Catalog) and play a crucial role in addressing the metadata crisis faced by enterprises as AI and data products become business-critical. You will lead the development of DataHub's Platform framework, which connects diverse data systems and powers our metadata collection capabilities. Your work will directly impact the deployment of AI systems at massive scale, affecting millions of users worldwide. This is a remote-friendly position with a range of benefits, including parental leave, medical, dental, and vision insurance, equity, and a monthly co-working space budget.
- Lead the development of DataHub's Platform framework, focusing on scalable and fault-tolerant ingestion systems for enterprise-scale metadata.
- Design and implement clean, intuitive APIs for the connector ecosystem, ensuring seamless integration with diverse data systems.
- Collaborate with engineering teams to address data discovery, lineage, and governance challenges, and contribute to the overall architecture of the metadata layer.
- Designed and scaled distributed systems
- Set up monitoring and alerting for services
- Strong systems knowledge and distributed architecture experience
- Proven track record solving complex technical challenges
- Advanced Python and API design expertise
- Experience with high-scale data processing or integration frameworks
- Designed indexing, storage, and data architectures to make large-scale data accessible to online services
- Hands-on experience developing in a tight loop with LLMs and applying best practices for scalable LLM development
- Built and maintained online applications serving live traffic at scale (100+ QPS)
- Python/TypeScript/Node.js - nice-to-have
- One of Java/Scala/Kotlin/C#/Go - very strong nice-to-have / borderline must-have
- CI/CD deployment pipelines
- AWS
- Kubernetes/Docker
- Microservice Architecture
- Early-stage startup experience
- Experience fine-tuning LLM-powered applications exposed to end users
- Experience building and maintaining services that make calls to LLMs in order to serve live traffic
- Open-source contributions
- Experience with DataHub or similar metadata/ETL frameworks (Airflow, Airbyte, dbt)
- 8+ years building production-grade distributed systems
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring