Skip to content
← Back to job listings

Digital Engineer

Sonata Software · Pune, Pune, Maharashtra, India

External listingfull-time22 days ago

About The Role

ROLE SUMMARY Data Engineers with hands-on expertise in PySpark and Fivetran to design, build, and maintain scalable data ingestion and transformation pipelines, at two levels: Junior (execution-focused, guided work) and Senior (ownership of architecture, performance, and mentoring). Both levels will work on integrating data from SaaS, database, and API sources into the cloud data warehouse, and building distributed processing jobs to support analytics and reporting. JUNIOR DATA ENGINEER — REQUIREMENTS • [Must-Have] 1–3 years of hands-on experience with PySpark (DataFrames, Spark SQL basics) — production or strong academic/project experience acceptable. • [Must-Have] Working knowledge of Fivetran or a similar ELT tool (Airbyte, Stitch) — setting up connectors, monitoring syncs. • [Must-Have] Solid SQL fundamentals and exposure to a cloud data warehouse (Snowflake / Redshift / BigQuery). • [Must-Have] Basic Python scripting ability for data tasks. • Exposure to at least one cloud platform (AWS / Azure / GCP), internship or project-level acceptable. • Eager to learn, works well under guidance from senior engineers; not expected to design architecture independently. NICE-TO-HAVE SKILLS (EITHER LEVEL) • Experience with Databricks and Delta Lake. • Familiarity with orchestration tools (Airflow, Databricks Workflows). • dbt experience for transformation/modeling. • Exposure to streaming technologies (Kafka, Spark Structured Streaming). • Relevant certifications (Databricks Certified Data Engineer, Fivetran certification, cloud data engineer certs). • Prior experience in [specific industry, e.g., healthcare, finance, retail], if relevant. RESPONSIBILITIES Junior: • Build and maintain PySpark transformations under senior engineer guidance. • Set up and monitor assigned Fivetran connectors; escalate sync issues. • Write and troubleshoot SQL queries for data validation. • Document pipeline steps and support operational runbooks. SCREENING / SOURCING NOTES FOR RECRUITER • For Senior: prioritize candidates who can clearly describe end-to-end pipeline ownership (source → Fivetran ingestion → PySpark transformation → warehouse load) and architecture decisions they made. • For Junior: focus on fundamentals — can they explain what a Spark DataFrame is, how a Fivetran sync works, and walk through a project (academic or professional) where they used both? • Ask for specific metrics from Senior candidates: data volumes handled, number of Fivetran connectors managed, pipeline runtime improvements. • Confirm recency: look for PySpark/Fivetran usage within the last [1–2] years, not just legacy ETL tools. • Target keywords for boolean search: "PySpark" AND "Fivetran" AND ("Snowflake" OR "Redshift" OR "BigQuery") AND ("Databricks" OR "EMR" OR "Glue"). • Add experience-range filters per level: Junior = 1–3 yrs total; Senior = 5+ yrs total, with 3+ yrs specifically in PySpark/Fivetran.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing