← Back to job listings
ME
Site Reliability / Infrastructure Engineer
Medal.tv · New York, United States
About The Role
Join Medal, a fast-growing startup in the video technology space. As a Site Reliability / Infrastructure Engineer, you will be responsible for ensuring the reliability and scalability of our infrastructure, which handles billions of clips and video ingestion pipelines. You will own the on-call rotation, drive postmortems, and work directly with engineering teams to meet their infrastructure needs. The ideal candidate has experience in startups, strong fluency in Terraform, and deep hands-on experience scaling and sharding relational databases.
- Ownership of the on-call rotation, driving postmortems, and working directly with engineering teams to meet their infrastructure needs.
- Managing and scaling infrastructure to handle billions of clips, video ingestion pipelines, and social features at a massive scale.
- Implementing and maintaining infrastructure-as-code using Terraform, and working with CI/CD tools like GitHub Actions in a production environment.
- The right person probably came through startups and scale-ups, has been in the room when things broke at 2am, has scaled databases under pressure, and knows the difference between a durable fix and a patch that buys you a week
- Infrastructure-as-code: Strong fluency in Terraform, with real experience owning infrastructure-as-code at scale
- Communication (crucial!): You flag issues clearly and rapidly during incidents and lead/write actionable postmortems
- CI/CD: You've worked with GitHub Actions in a production environment
- Experience at startups: You are comfortable in an environment of rapid growth where scaling up is a priority
- Elasticsearch depth: Hands-on experience running ES for user-facing features, not just as a log sink
- Great judgment: You know the difference between a durable, sustainable fix and a patch that buys you a week
- Incident response instincts: You can work a P0 calmly, communicate clearly under pressure, and run a postmortem that prevents recurrence
- Database scaling: Deep, hands-on experience scaling and sharding relational databases (MySQL, Postgres) in production
- GCP depth: You know it maybe a little too well: Kubernetes, VPC, IAM, Cloud Logging, and the managed services ecosystem
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring