Skip to content
← Back to job listings

Senior Site Reliability Engineer

ServiceTitan · United States

External listingfull-time5 days ago

About The Role

Join ServiceTitan as a Senior Site Reliability Engineer. In this role, you will be responsible for the reliability and health of our cloud-based applications, designing signals to detect issues and building systems to improve performance as we scale. You will participate in an on-call rotation, design and maintain observability dashboards, operate and improve our Kubernetes-based compute platform, and collaborate with product engineering teams. This position offers flexible time off, fully paid medical, dental, and vision coverage, and various other benefits.

  • Participate in an on-call rotation, using runbooks and playbooks to diagnose and resolve production issues.
  • Design, build, and maintain observability dashboards and alerting grounded in Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
  • Operate and improve our Kubernetes-based compute platform, which runs the large majority of our infrastructure.
  • SRE principles: practical experience with SLIs, SLOs, and error budgets — able to speak to how you've defined and monitored these on real systems, not just definitions
  • Kubernetes (must-have): strong, hands-on understanding of Kubernetes as a system
  • Cloud engineering & networking: solid grounding in AWS or Azure, including networking fundamentals (subnetting, IP addressing)
  • 8-10+ years of relevant hands-on experience
  • Strong programming skills with the ability to build web applications — ideally with solid working knowledge of .NET and <ASP.NET>. We're also open to strong Python (Flask, FastAPI) or Java (Spring) backgrounds. The coding assessment will be tailored to whichever language/framework you're most comfortable in
  • Observability: deep experience with at least one modern observability stack (OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch) and the ability to translate that understanding across tools
  • Strong production troubleshooting skills — comfortable diagnosing issues under pressure
  • CI/CD: strong understanding of a CI/CD system — GitHub Actions preferred, but TeamCity, Azure DevOps, or GitLab CI experience is acceptable
  • Experience with distributed systems and their common failure modes (retries, timeouts, cascading failures)
  • Nice-to-have: database experience (not mandatory — databases are monitored by the same team, not owned individually)
  • You're comfortable guiding and making decisions with limited information, and capable of operating within the trade-offs between solving for immediate needs versus bigger-scale solutions
  • You feel rewarded by developing an operability culture in a quickly growing and changing environment, and you're comfortable owning a wide and diverse set of problem areas
  • You're someone who enjoys being directly accountable for the reliability of a business-critical, large-scale enterprise system
  • Being human isn’t about checking every box on a list. It’s about the experiences we have, people we meet, and the perspectives we share. So, if you have the skills but are hesitant to apply because of your background, apply anyway

This is an external listing. JobSpring does not represent or verify the employer. Report this listing