← Back to job listings
TA
Customer Success Engineer (GPU Cluster)
Together AI · San Francisco, United States
About The Role
Join Together AI as a Customer Success Engineer, where you will be the primary technical point of contact for a strategic customer relationship. You will drive structured engagement, manage issue lifecycles, and maintain deep technical expertise across the customer's infrastructure stack. This role requires a strong ownership mindset, proficiency in infrastructure monitoring and observability tooling, and experience with DC operations and facilities coordination. Enjoy competitive health insurance, flexible time off, and team-driven celebrations and events.
- Servir como el punto de contacto técnico principal para un cliente estratégico, gestionando la relación técnica de extremo a extremo en todas las áreas de infraestructura.
- Dirigir el compromiso estructurado a través de cadencias regulares, incluyendo informes de estado, reuniones de dirección técnica y revisiones comerciales ejecutivas.
- Gestionar el ciclo de vida de los problemas, la escalación y la autoría de RCA en todas las áreas de infraestructura en asociación con los equipos de soporte, SRE, DC Ops y ingeniería.
- Hands-on experience with large-scale Ethernet and InfiniBand fabric architecture
- Strong ownership mindset for incident management, RCA authorship, and executive-level customer communication
- 5+ years in a customer-facing technical role, with 2+ years in dedicated technical account management or solutions architecture for large-scale AI or HPC infrastructure
- Working knowledge of enterprise storage systems, including high-density NVMe, parallel file systems, and metadata infrastructure
- Proficiency in infrastructure monitoring and observability tooling (Prometheus, Grafana, or equivalent)
- Deep expertise in GPU infrastructure — GPU health diagnostics, RMA workflows, and hardware acceptance testing
- Experience with DC operations, facilities coordination, and hosting provider SLA management
- Proven ability to manage multiple concurrent workstreams with hyperscaler-level rigor and communication standards
- Proficiency in Python, Bash, or infrastructure automation tools preferred
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring