Skip to content
← Back to job listings

Customer Support Engineer (Inference)

Together AI · San Francisco, United States

External listingfull-time2 months ago

About The Role

Join Together AI, a pioneering AI company, as a Customer Support Engineer. In this role, you will be the first line of defense for customers building training, fine-tuning, and inference solutions. You will tackle complex technical challenges, collaborate with various teams, and drive continuous improvement of our offerings. This is an exciting opportunity for a deeply technical professional passionate about AI and customer success.

  • Engage directly with customers to tackle and resolve complex technical challenges involving our cutting-edge GPU clusters and our inference and fine-tuning services.
  • Collaborate seamlessly across Engineering, Research, and Product teams to address customer concerns; collaborate with senior leaders both internally and externally to ensure the highest levels of customer satisfaction.
  • Transform customer insights into action by identifying patterns in support cases and working with Engineering and Go-To-Market teams to drive Together’s roadmap.
  • Ability to work cross-functionally with teams such as Sales, Engineering, Support, Product and Research to drive customer success
  • Complex technical problem solving and troubleshooting, with a proactive approach to issue resolution
  • Strong technical background, with knowledge of AI, ML, GPU technologies and their integration into high-performance computing (HPC) environments
  • Familiarity with operating storage systems in HPC environments such as Vast and Weka
  • 5+ years of experience in a customer-facing technical role with at least 1 year in a support function in AI
  • Strong knowledge of Python, TypeScript, and/or JavaScript with testing/debugging experience using curl and Postman-like tools
  • Strong sense of ownership and willingness to learn new skills to ensure both team and customer success
  • Familiarity with infrastructure services (e.g., Kubernetes, SLURM), infrastructure as code solutions (e.g., Ansible) high-performance network fabrics, NFS-based storage management, container infrastructure, and scripting and programming languages
  • Familiarity with inspecting and resolving network-related errors
  • Ability to operate in dynamic environments, adept at managing multiple projects, and comfortable with frequent context switching and prioritization
  • Excellent communication and interpersonal skills, with the ability to explain complex technical concepts to non-technical stakeholders
  • Foundational understanding in the installation, configuration, administration, troubleshooting, and securing of compute clusters

This is an external listing. JobSpring does not represent or verify the employer. Report this listing