← Back to job listings
PD
Senior Software Engineer (Data Acquisition)
People Data Labs · United States
About The Role
Join People Data Labs (PDL), a leading provider of people and company data. As a Senior Software Engineer (Data Acquisition), you will play a crucial role in building standalone data products and improving existing datasets. You will work with a team to enhance our data acquisition and processing platform, develop web crawling technologies, and design backend services. This remote-first position offers unlimited PTO, monthly employee stipends, and comprehensive medical, dental, and vision insurance.
- Contribuer à l'architecture et à l'amélioration de notre plateforme d'acquisition et de traitement des données, en augmentant la fiabilité, le débit et l'observabilité.
- Utiliser et développer des technologies de web crawling pour capturer et cataloguer des données sur Internet.
- Construire, exploiter et faire évoluer des systèmes distribués à grande échelle qui collectent, traitent et livrent des données provenant du web.
- Familiarity with network architecture and debugging (HTTP, DNS, proxies, packet capture and analysis)
- Strong grasp of software architecture and backend fundamentals; you can reason clearly about concurrency, scalability, and fault tolerance
- Familiarity with message queues, orchestration, and distributed task systems (Kafka, SQS, Airflow, etc.)
- Solid understanding of browser rendering pipeline, web application architecture (auth, cookies, http request / response)
- Experience designing or maintaining resilient data ingestion, API integration, or ETL systems
- 7+ years of professional experience building or operating backend or infrastructure systems at scale
- Proficiency with Linux / Unix command-line tools and system resource management
- Experience evaluating and monitoring data quality, ensuring consistency, completeness, and reliability across releases
- Solid understanding of distributed systems concepts: parallelism, asynchronous programming, backpressure, and message-driven design
- Solid programming experience in Python, Go, Rust, or similar, including experience with async / await, coroutines, or concurrency frameworks
- Write and maintain technical design documents, including pipeline design, schema design, and data flow diagrams
- Balance pragmatism with craftsmanship, shipping reliable systems while continuously improving them
- Scope and break down complex projects into deliverable milestones, and communicate progress, risks, and blockers effectively
- Communicate clearly and thoughtfully in writing (Slack, docs, design proposals)
- Work independently in a fast-paced, remote-first environment, proactively unblocking themselves and collaborating asynchronously
- Degree in a quantitative field such as computer science, mathematics, or engineering
- Experience as a Red Teamer
- Experience working on large-scale data ingestion, crawling, or indexing systems
- Experience with Apache Spark, Databricks, or other distributed data platforms
- Experience with streaming data systems (Kafka, Pub/Sub, Spark Streaming, etc.)
- Proficiency with SQL and data warehousing (Snowflake, Redshift, BigQuery, or similar)
- Experience with cloud platforms (AWS preferred, GCP or Azure also great)
- Understanding of modern data storage and design patterns (parquet, Delta Lake, partitioning, incremental updates)
- Knowledge of modern data design and storage patterns (e.g., incremental updating, partitioning and segmentation, rebuilds and backfills)
- Experience building and maintaining data pipelines on modern big-data or cloud platforms (Databricks, Spark, or equivalent)
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring