← Back to job listings
GT
Senior Platform Reliability Engineer
Grow Therapy · United States
About The Role
Join Grow as a Senior Platform Reliability Engineer, where you'll define and scale reliability as a core capability. You'll work across the organization to establish standards for observability, SLOs/SLAs, and incident response, while also creating self-service tooling to facilitate adoption. This high-impact role will allow you to drive cultural and technical change, enabling teams to independently build and operate reliable systems at scale.
- Defining and establishing reliability standards, frameworks for SLOs/SLAs, and operational readiness.
- Improving observability and measurement by identifying gaps in metrics, logging, and tracing.
- Driving adoption of reliability practices across teams through education, influence, and guidance.
- Deep understanding of reliability principles: You’ve defined or worked with SLOs/SLAs, understand error budgets, and have experience improving reliability through measurement and iteration
- Systems thinker: You’re able to zoom out, identify patterns across teams and services, and design solutions that scale beyond a single system
- Strong communicator and influencer: You can drive change across teams without direct authority, balancing pragmatism with long-term vision
- Experienced in production systems: You have 6+ years of experience operating and improving reliability of production systems at scale
- Observability expertise: You’ve worked with modern observability tooling (we use DataDog) and understand how to build actionable monitoring systems across metrics, logs, and traces
- Impact-oriented: You focus on outcomes over output and care deeply about improving real reliability outcomes—not just adding processes
- Team player: You collaborate well, communicate with empathy, and enjoy mentoring and learning from others
- Strong foundation in cloud and infrastructure: You have hands-on experience with AWS, Kubernetes (e.g., EKS), and infrastructure as code tools like Terraform
- Self-directed: You thrive in ambiguous environments and are comfortable defining problems, proposing solutions, and executing independently
- You’ve worked with both SaaS (e.g., DataDog) and self-managed observability stacks
- You’ve built internal tooling or platforms used by multiple teams
- If you’re excited about this role but don’t check every box, we encourage you to apply. At Grow, we value diverse experiences, transferable skills, and the unique strengths each person brings
- You have experience with database reliability and performance (we use PostgreSQL)
- You have experience designing service-level scorecards or compliance/reporting systems
- You were previously a product engineer and bring empathy for developer experience
- You’ve helped introduce or scale reliability practices in a growing organization
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring