Skip to content
← Back to job listings

Senior Machine Learning Engineer (Voice AI)

Together AI · San Francisco, United States

External listingfull-time23 days ago

About The Role

Join Together AI, a company dedicated to building the best inference infrastructure for voice applications. As a Senior Machine Learning Engineer, you will drive the model serving layer for voice workloads, optimizing how we serve models like Whisper, Parakeet, Orpheus, and Kokoro. You will work hands-on with inference engines, profile GPU utilization, design batching strategies for streaming audio, and ensure new model architectures can go from research to production quickly. This is a foundational hire on a small, high-impact team, and you will have the opportunity to shape how Together serves voice models as the industry evolves.

  • Conduire la couche de service de modèle pour les charges de travail vocales, en travaillant directement avec des moteurs d'inférence tels que TRT-LLM et SGLang.
  • Optimiser la performance d'inférence pour les modèles vocaux (STT, TTS, speech-to-speech), en visant des temps de réponse et un débit de premier ordre.
  • Collaborer avec des partenaires de modèles pour intégrer et optimiser leurs modèles fonctionnant sur l'infrastructure de Together.
  • 5+ years of experience in ML engineering, with a focus on model serving, inference optimization, or ML infrastructure
  • Experience with speech and audio ML (ASR, TTS architectures, audio signal processing) is a strong plus but not required — you can learn this quickly if you have strong ML engineering fundamentals
  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field, or equivalent practical experience
  • Strong proficiency in Python and PyTorch; experience with GPU profiling and optimization (CUDA, memory management, kernel-level debugging)
  • Track record of shipping ML systems to production with measurable performance improvements
  • Familiarity with audio codecs and tokenization schemes (SNAC, Encodec, DAC) is a plus
  • Experience training or fine-tuning speech models is a plus
  • Hands-on experience with LLM serving engines (vLLM, SGLang, TensorRT-LLM, or similar) — comfortable reading and modifying engine internals, not just using APIs
  • Strong product sense — you think about what developers building voice apps actually need, not just what's technically interesting
  • Comfort working on a small, early-stage team where you'll wear multiple hats and move fast

This is an external listing. JobSpring does not represent or verify the employer. Report this listing