Skip to content
← Back to job listings

Senior Platform Reliability Engineer

Grow Therapy · United States

External listingfull-timeabout 1 month ago

About The Role

Join Grow as a Senior Platform Reliability Engineer, where you'll define and scale reliability as a core capability. You'll work across the organization to establish standards for observability, SLOs/SLAs, and incident response, while also creating self-service tooling to facilitate adoption. This high-impact role will allow you to drive cultural and technical change, enabling teams to independently build and operate reliable systems at scale.

  • Defining and establishing reliability standards, frameworks for SLOs/SLAs, and operational readiness.
  • Improving observability and measurement by identifying gaps in metrics, logging, and tracing.
  • Driving adoption of reliability practices across teams through education, influence, and guidance.
  • Deep understanding of reliability principles: You’ve defined or worked with SLOs/SLAs, understand error budgets, and have experience improving reliability through measurement and iteration
  • Systems thinker: You’re able to zoom out, identify patterns across teams and services, and design solutions that scale beyond a single system
  • Strong communicator and influencer: You can drive change across teams without direct authority, balancing pragmatism with long-term vision
  • Experienced in production systems: You have 6+ years of experience operating and improving reliability of production systems at scale
  • Observability expertise: You’ve worked with modern observability tooling (we use DataDog) and understand how to build actionable monitoring systems across metrics, logs, and traces
  • Impact-oriented: You focus on outcomes over output and care deeply about improving real reliability outcomes—not just adding processes
  • Team player: You collaborate well, communicate with empathy, and enjoy mentoring and learning from others
  • Strong foundation in cloud and infrastructure: You have hands-on experience with AWS, Kubernetes (e.g., EKS), and infrastructure as code tools like Terraform
  • Self-directed: You thrive in ambiguous environments and are comfortable defining problems, proposing solutions, and executing independently
  • You’ve worked with both SaaS (e.g., DataDog) and self-managed observability stacks
  • You’ve built internal tooling or platforms used by multiple teams
  • If you’re excited about this role but don’t check every box, we encourage you to apply. At Grow, we value diverse experiences, transferable skills, and the unique strengths each person brings
  • You have experience with database reliability and performance (we use PostgreSQL)
  • You have experience designing service-level scorecards or compliance/reporting systems
  • You were previously a product engineer and bring empathy for developer experience
  • You’ve helped introduce or scale reliability practices in a growing organization

This is an external listing. JobSpring does not represent or verify the employer. Report this listing