Skip to content
← Back to job listings

Customer Success Engineer (GPU Cluster)

Together AI · San Francisco, United States

External listingfull-time22 days ago

About The Role

Join Together AI as a Customer Success Engineer, where you will be the primary technical point of contact for a strategic customer relationship. You will drive structured engagement, manage issue lifecycles, and maintain deep technical expertise across the customer's infrastructure stack. This role requires a strong ownership mindset, proficiency in infrastructure monitoring and observability tooling, and experience with DC operations and facilities coordination. Enjoy competitive health insurance, flexible time off, and team-driven celebrations and events.

  • Servir como el punto de contacto técnico principal para un cliente estratégico, gestionando la relación técnica de extremo a extremo en todas las áreas de infraestructura.
  • Dirigir el compromiso estructurado a través de cadencias regulares, incluyendo informes de estado, reuniones de dirección técnica y revisiones comerciales ejecutivas.
  • Gestionar el ciclo de vida de los problemas, la escalación y la autoría de RCA en todas las áreas de infraestructura en asociación con los equipos de soporte, SRE, DC Ops y ingeniería.
  • Hands-on experience with large-scale Ethernet and InfiniBand fabric architecture
  • Strong ownership mindset for incident management, RCA authorship, and executive-level customer communication
  • 5+ years in a customer-facing technical role, with 2+ years in dedicated technical account management or solutions architecture for large-scale AI or HPC infrastructure
  • Working knowledge of enterprise storage systems, including high-density NVMe, parallel file systems, and metadata infrastructure
  • Proficiency in infrastructure monitoring and observability tooling (Prometheus, Grafana, or equivalent)
  • Deep expertise in GPU infrastructure — GPU health diagnostics, RMA workflows, and hardware acceptance testing
  • Experience with DC operations, facilities coordination, and hosting provider SLA management
  • Proven ability to manage multiple concurrent workstreams with hyperscaler-level rigor and communication standards
  • Proficiency in Python, Bash, or infrastructure automation tools preferred

This is an external listing. JobSpring does not represent or verify the employer. Report this listing