Skip to content
← Back to job listings

Software Engineer (Data Infrastructure)

Peregrine · San Francisco, United States

External listingfull-timeabout 2 months ago

About The Role

Join Peregrine, a fast-growing company focused on data infrastructure. As a Data Infrastructure Engineer, you will have deep ownership over the data layer that underpins everything we do. You will architect and build systems that ingest, store, and serve massive volumes of real-time operational data. This is an individual contributor role for someone who thrives on hard technical problems and brings the experience and judgment to shape foundational infrastructure decisions.

  • Architecting and building systems for ingesting, storing, and serving massive volumes of real-time operational data.
  • Designing and operating a high-throughput, real-time data integration platform across diverse customer environments.
  • Building and optimizing distributed data processing pipelines with Apache Spark and adjacent streaming technologies.
  • Experience with Kubernetes and containerized deployment of data workloads
  • Located in San Francisco and open to working in office
  • Extensive hands-on experience with Apache Spark for batch and streaming data processing at scale
  • 2-5 years of experience operating large-scale data infrastructure systems in production environments
  • Experience with open table formats, particularly Apache Iceberg — including schema evolution, partitioning strategies, compaction, and time travel
  • Strong technical vision with the ability to translate complex data requirements into clean, durable infrastructure designs
  • Thrive on ambiguity and are energized by defining the right solution to hard, open-ended problems
  • Degree in Computer Science, Engineering, or a related field, or equivalent practical experience
  • Desire to own significant portions of the data stack end-to-end, from ingestion to serving
  • Deep passion for data infrastructure — you care about building systems that are correct, fast, and resilient at scale
  • Experience with AWS or comparable cloud platforms, including S3-based data lake architectures
  • Experience with data pipeline orchestration using Airflow or similar tools
  • Strong software engineering fundamentals in Python and/or Scala, with a track record of writing production-quality code
  • Committed to operational excellence — you build things you’re proud to operate
  • Background in real-time data integration and stream processing, leveraging technologies such as Apache Kafka, Apache Flink, or equivalents

This is an external listing. JobSpring does not represent or verify the employer. Report this listing