Skip to content
← Back to job listings

Data Platform Engineer

Yotta Infrastructure · Delhi, Delhi, India

External listingFull-timeabout 2 months ago

About The Role

About Yotta Yotta Data Services is India’s leading sovereign AI infrastructure, cloud platform and data centre services company, enabling enterprises, governments, startups, and digital platforms to build, deploy, and scale next-generation AI and digital workloads securely within India. With hyperscale data centre campuses in Navi Mumbai and Greater Noida (Delhi NCR), advanced GPU-powered AI infrastructure, and a comprehensive ecosystem of cloud, AI, hosting, cybersecurity, and managed platform services, Yotta delivers high-performance, scalable, and compliant digital infrastructure built for the AI era. Yotta is at the forefront of powering India’s sovereign AI and digital transformation journey through world-class infrastructure, deep technology partnerships, and fully India-hosted enterprise-grade platforms. Job Scope We are looking for an enthusiastic Junior Data Platform Engineer to support and manage our Apache Spark, Apache Airflow, and JupyterHub environments. This role is ideal for someone with a strong foundation in Python and Linux, who is eager to build a career in big data engineering and data platform administration. You will work closely with senior engineers to ensure smooth operation, deployment, and optimization of our data processing ecosystem. Total /Relevant Experience · 1+ years experience Key Responsibilities · Assist in the setup, monitoring, and maintenance of Apache Spark clusters, Apache Hive, Hadoop and Airflow environments. · Support the development and scheduling of data pipelines using Airflow DAGs and Python scripts. · Help manage and configure JupyterHub for multi-user access and integration with Spark. · Monitor cluster health and performance under guidance and assist in troubleshooting Spark job failures. · Write and maintain Python automation scripts for data workflows, ETL, and process automation. · Participate in code reviews, documentation, and deployment activities. · Learn and follow best practices for distributed data processing, CI/CD, and DevOps workflows. · Collaborate with senior engineers and data scientists to implement improvements and new features. Must-have skill · Basic understanding of Apache Spark, Apache Hive and Hadoop File System. · Familiarity with Apache Airflow (understanding of DAGs, scheduling, and task dependencies). · Hands-on experience with Python scripting (data processing, automation, or API interaction). · Comfortable working in Linux and container environments (command line, system logs, process management). · Good understanding of data processing concepts, including ETL and distributed computing. · Basic knowledge of Git and version control. Good-to-Have Skills · Exposure to Jupyter / JupyterHub for collaborative notebook environments. · Knowledge of Docker or Kubernetes. · Knowledge of Hadoop and Apache Spark Cluster. · Familiarity with SQL and working with structured/unstructured data. · Experience with cloud platforms (AWS, GCP, or Azure) is a plus. · Interest in big data (Hadoop), DevOps, and data pipeline automation. Qualifications Criteria · Bachelor’s or any relevant Degree. Certification Criteria · NA Shifts Timing, (If rotational, please specify) · General Shift Number of Interview Rounds Name of interviewer · 3 rounds Behavioral Attributes: Art of skilful conversation Creativity & Problem Solving Learning on fly Business Acumen Building Trust Customer Focus Intellectual Horsepower (Functional Skills) Action Orientation & Accountability Process-Quality Excellence Prioritizing, Planning & organizing Listening, Sensing, Observing Developing Direct Reports Peripheral Building Collaborative Relationships Company Values Customer Centricity Agility Integrity Innovation Trust and Transparency Happiness for all

This is an external listing. JobSpring does not represent or verify the employer. Report this listing