Senior Data Engineer
beigene · 大连市, 辽宁, 中国
About The Role
General Description The Senior Data Engineer in the GTS – Data & Analytics team will be based in Dalian, China as part of our shared service center strategy and will help build a modern, scalable data engineering capability that serves global business needs. This role is responsible for designing, building, and operating scalable, reliable, and compliant data pipelines, data products, and analytical data products on the Databricks lakehouse platform to support enterprise analytics, AI, and decision support. This role partners closely with analytics engineers, data scientists, platform teams, and business stakeholders to translate business requirements into resilient, reusable, and governed data assets. The ideal candidate combines strong hands-on engineering depth in Databricks, Spark, Python, and SQL with a practical product mindset and the ability to deliver reliable, consumption-ready data solutions in collaboration with global teams. The position plays a critical role in ensuring data consistency, availability, performance, and security across platforms, with a strong focus on operational excellence and data governance. We are building an AI-proficient data organization, and this role is expected to use AI as a core component of the software development lifecycle—from requirements acceleration and solution design to code generation, testing, documentation, troubleshooting, and continuous improvement—while maintaining strong standards for quality, security, and compliance. Essential Functions of the Job Design, develop, and maintain end-to-end batch and/or streaming data pipelines on the Databricks lakehouse platform to deliver trusted, analytics-ready data products aligned to medallion architecture. Integrate data from multiple source systems, translating business requirements into resilient, reusable technical solutions with clear data contracts and schema governance. Build and manage scalable data models and curated data layers following medallion architecture patterns (bronze/silver/gold) to ensure consistent business definitions, metrics, and governed consumption by downstream BI, analytics, and AI/BI use cases. Develop and maintain data solutions and interfaces for downstream use cases including reporting, self-service analytics, and emerging conversational analytics / AI/BI experiences, ensuring data availability, stability, and controlled evolution of schemas and metrics. Operate and support production data services, including monitoring, alerting, incident resolution, root cause analysis, and release management. Ensure data quality by implementing automated validation, monitoring, reconciliation controls, and clearly defined data SLAs with structured issue resolution processes. Implement and enforce data governance requirements, including metadata management via Unity Catalog, data lineage, discovery, access control, approval workflows, and auditability. Optimize data pipeline performance, reliability, and cost through efficient Spark/PySpark data processing, Delta Lake storage design, partitioning strategies, and query tuning. Conduct design and code reviews, providing technical guidance and mentoring to junior engineers and external resources. Produce and maintain clear technical documentation, including architecture diagrams, data flows, data contracts, semantic documentation, and operational runbooks. Build and maintain high-quality, governed data foundations that support downstream conversational analytics, AI/BI experiences, and natural language querying through well-structured metadata, business definitions, and semantic context. Leverage AI-powered tools (e.g., code assistants, automated testing, documentation and troubleshooting tools) as a core part of the development lifecycle to improve productivity, code quality, and solution reliability. Supervisory Responsibilities This role does not typically have direct people management responsibilities but provides technical leadership and guidance to junior engineers and external or project-based resources as required. Computer Skills Advanced SQL and strong proficiency in Python; PySpark experience required, with Scala as a plus Strong hands-on experience with distributed data processing frameworks, especially Apache Spark and PySpark, for large-scale batch and streaming workloads Strong understanding of modern data warehouse and lakehouse architecture, including medallion design patterns and the delivery of reusable data products and analytical data products Deep hands-on experience with Databricks for enterprise data engineering, including Delta Lake, Databricks SQL, notebooks, workflows/jobs, repos, cluster or compute management, and production pipeline development; strong familiarity with Unity Catalog for centralized metadata management, fine-grained access control, data lineage, discovery, and permissions governance across workspaces Familiarity with Databricks AI/BI capabilities, including Genie Code and Genie Spaces, and the ability to help create governed, high-quality data foundations that support natural language and business-facing analytical experiences Hands-on experience with Microsoft Azure data services, including Azure Data Factory, ADLS Gen2/Blob Storage, Key Vault, and strong understanding of Azure security and access control models such as RBAC and managed identities to support data permissions governance and compliance Proficiency with version control and CI/CD tools (e.g., Databricks bundles, Git, Azure DevOps, GitHub) Understanding of cloud data platforms and security concepts, including role-based access control Experience with data quality, monitoring, and observability tools Ability to document technical solutions using standard documentation and diagramming tools Other Qualifications 5+ years of experience in data engineering, software engineering, or related technical roles Experience leveraging AI-powered tools (e.g., code assistants, automated testing, documentation or troubleshooting tools) to improve development efficiency, code quality, and ongoing maintenance of data engineering projects Strong communication skills and ability to work effectively with cross-functional stakeholders Experience working in Pharma, BioTech or large-scale corporate environments is a plus Databricks certification is preferred, especially [Databricks Certified Data Engineer Associate]() or higher; Azure certification or other relevant modern data engineering credentials are also valued Good written and spoken English communication skills are required, with the ability to collaborate clearly and confidently in global technical discussions and produce high-quality technical documentation. Travel Occasional travel may be required, up to 10%, depending on business needs. BeOne Global Competencies Customer Focus Collaboration Accountability Integrity & Trust Results Orientation Innovation Continuous Improvement Data-Driven Decision Making Effective Communication Inclusion & Respect Strategic Thinking Learning Agility 百济神州全球胜任力 当我们通过以下十二项全球胜任力,展现出 "患者为先"、"无界协作"、"锐意创新 "和 "追求卓越 "的价值观时,我们就能帮助全世界更多患者获得更多负担得起的药品。 ●团队协作 ●提供并征求坦诚及可行的反馈 ●自我认知 ●兼容并蓄 ●积极主动 ●开拓精神 ●持续学习 ●拥抱变化 ●结果导向 ●分析性思维/数据分析 ●卓越财务 ●清晰沟通 BeOne Global Competencies When we exhibit our values of Patients First, Collaborative Spirit, Bold Ingenuity and Driving Excellence, through our twelve global competencies below, we help get more affordable medicines to more patients around the world. ●Fosters Teamwork ●Provides and Solicits Honest and Actionable Feedback ●Self-Awareness ●Acts Inclusively ●Demonstrates Initiative ●Entrepreneurial Mindset ●Continuous Learning ●Embraces Change ●Results-Oriented ●Analytical Thinking/Data Analysis ●Financial Excellence ●Communicates with Clarity 求职者隐私申明: 百济神州致力于尊重和保护您的个人信息权利,并承诺依据合法、正当、必要和诚信的原则处理您的个人信息(包括个人敏感信息 )。 由于百济神州在全球范围内开展业务,我们可能需要基于人力资源管理等合理业务目的而将您的个人信息发送和/或存储在位于您所在国家以外其他国家(例如:美国)的服务器和数据库中,详情参见百济神州《求职者隐私政策》(百济神州官网 - 隐私政策 - 求职者隐私政策)。 如您主动向我们提供您的简历信息或其他个人信息,则视为您已经充分理解并确认接受百济神州《求职者隐私政策》内容。如您对此有任何疑问的,请勿提交简历信息或其他个人信息。 BeOne is committed to respect and protect your personal information rights, and will process your personal information, including your sensitive personal information, based on the principles of legality, legitimacy, necessity, and integrity. Due to the reasonable business need for human resource management as a result of BeOne’s global operation, your personal information may be transferred and/ or stored in a server/database located in a third country (e.g., the United States) other than your own country. For further details, please refer to BeOne Job Applicant Privacy Policy (BeOne official website - Privacy Policy - Job Applicant Privacy Policy). If you voluntarily provide your resume or other personal information to us, it is deemed as you have thoroughly acknowledged and accepted BeOne Job Applicant Privacy Policy. If you have any concern, please DO NOT submit your resume or any other personal information.
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring