Skip to content
← Back to job listings

Staff Engineer (Site Reliability)

Babylist · Ottawa, Canada

External listingfull-timeabout 1 month ago

About The Role

Join Babylist, a rapidly growing company that started as an e-commerce and registry platform and is now expanding into health, media, mobile, and new product surfaces. As a Staff Engineer (Site Reliability), you will be at the center of keeping our platform reliable, fast, and scalable. You will own the infrastructure and reliability practices that support over 9 million users and the engineers who build for them. This is a staff-level role with real cross-team visibility, and your work will have a significant impact across the entire product organization.

  • Infrastructure ownership — manage and evolve our AWS environment using Terraform, keeping EKS clusters, databases, and core services current and performant.
  • CI/CD reliability — own the speed and reliability of our CI systems for the full Engineering org — every deploy starts here.
  • Monitoring & alerting standards — establish and socialize best practices so the right people get paged for the right reasons.
  • You naturally reach for AI in your work — at Babylist, every team uses AI daily. You're already using it to move faster and improve your output, and you stay curious about what's coming next
  • Deep hands-on Terraform expertise — you own IaC, not just contribute to it
  • Strong observability instincts — Datadog, Sentry, PagerDuty, Cronitor — you build alerting that's actionable, not noisy
  • Experienced with on-call and incident management — you've run the post-mortems and actually changed things afterward
  • Comfortable designing and improving CI/CD systems — CircleCI, GitHub Actions, or similar; you care about developer velocity, not just pipeline uptime
  • Proven AWS experience at scale — EKS, RDS, cloud networking, DNS, CDNs, load balancers — you know the gotchas
  • Comfortable supporting developers across local, staging, and production — you're a resource, not a gatekeeper
  • Experienced operating Kubernetes in production — you've debugged the hard stuff, not just deployed the easy stuff

This is an external listing. JobSpring does not represent or verify the employer. Report this listing