← Back to job listings
BA
Site Reliability Engineer
Baseten · United States
About The Role
Join Baseten as a Site Reliability Engineer, where you'll be the primary technical owner for our strategic customers, ensuring smooth deployment, performance, and reliability of ML workloads in production. You'll diagnose and resolve runtime issues, debug infrastructure issues, lead incident response during outages, and serve as the technical owner for top enterprise accounts. This remote-first position offers unlimited PTO, full healthcare coverage, paid parental leave, and a learning and development budget.
- Assurer le succès technique et l'issue à long terme des comptes stratégiques et d'entreprise en gérant et en résolvant les escalades.
- Diagnostiquer et résoudre les problèmes d'exécution liés à la latence, au comportement de la mémoire, à l'utilisation du GPU, à la concurrence et à la gestion du cycle de vie du modèle.
- Diriger la réponse aux incidents lors des pannes ou des escalades, en gérant la coordination entre le produit, l'ingénierie et les ventes.
- Deep Kubernetes troubleshooting expertise, including advanced resource debugging, pod/runtime analysis, and log-based diagnostics using observability tooling such as Grafana, Loki, and Prometheus
- 3+ years of experience in a fast-paced, high-growth, or customer-facing engineering environment
- Strong communication skills and executive presence during high-visibility situations, ensuring technical clarity and customer confidence
- Ability to translate recurring technical pain points into roadmap-level insights, documentation improvements, or product enhancements
- Strong infrastructure debugging ability across container orchestration, networking, and service dependencies, with hands-on experience supporting production-grade clusters
- Experience managing high-severity incidents with major customers, including SLAs, post-incident reviews, and clear communication throughout escalations
- Proven project management and organizational skills with an ownership mindset, able to manage multiple complex, multi-stakeholder initiatives in parallel — including issue resolution, root-cause analysis, and feature delivery
- Familiarity with running high-performance AI models and workloads, including troubleshooting ML pipelines from preprocessing through inference and serving
- Experience implementing or managing ticketing and incident-response systems such as Zendesk or Pylon
- Familiarity with Helm, Flux, CI/CD tooling, or scripting automations to improve deployment, release, or operational workflows
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring