Skip to content
← Back to job listings

Site Reliability Engineer (SRE)

Mithril · Palo Alto, United States

External listingfull-time22 days ago

About The Role

Join Mithril, a cutting-edge GPU orchestration platform, as a Site Reliability Engineer (SRE). In this role, you will be a key contributor to the stability and performance of our global platform, building automation, observability, and tooling to ensure fast and reliable access to infrastructure. You will work directly with our founding team on high-impact infrastructure decisions and have real ownership of your work. The primary focus of this role will be platform reliability and infrastructure automation, with opportunities to shape the future of our marketplace and company.

  • Contribuer à la stabilité et à la performance de la plateforme d'orchestration GPU de Mithril.
  • Construire l'automatisation, l'observabilité et les outils qui permettent à Mithril de coordonner des calculs avancés à grande échelle.
  • Participer à la conception des SLO et à l'orchestration de la capacité, en ayant une visibilité claire sur les dynamiques techniques et commerciales de l'entreprise.
  • The profile we're hiring for combines strong systems instincts with the ability to build and own tooling end-to-end. Candidates should be able to point to infrastructure or automation they've built that is still running in production
  • Cloud proficiency in at least one major provider (AWS, GCP, or Azure), including practical understanding of cloud networking fundamentals (VPC, DNS, load balancing, security groups)
  • Disciplined troubleshooter: calm under pressure during production incidents, with a rigorous approach to RCA and long-term remediation
  • Clear communicator: able to document processes and explain technical trade-offs to engineering and non-engineering teammates
  • 3+ years of experience in SRE, Production Engineering, or Infrastructure roles at a high-growth technology company
  • Coding ability: proficiency in Python or equivalent (Go, Rust, etc.) — you build tools and services, not just scripts. Willing to pick up new languages as needed
  • Linux fundamentals: strong command of Linux systems, TCP/IP networking, and security best practices
  • Hands-on Kubernetes experience: comfortable managing clusters, deployments, and troubleshooting production incidents in a multi-tenant environment
  • Experience with GPU/TPU-accelerated workloads or AI/ML infrastructure
  • Exposure to multi-cloud deployments or niche/specialized cloud providers (e.g., CoreWeave, Lambda Labs, Nebius)
  • Familiarity with distributed systems concepts — service discovery, circuit breakers, consensus protocols
  • Production experience with Prometheus, Grafana, or OpenTelemetry
  • Prior experience in a high-growth startup environment where infrastructure scope expands faster than headcount

This is an external listing. JobSpring does not represent or verify the employer. Report this listing