Skip to content
← Back to job listings

Staff Machine Learning Engineer (Voice AI)

Together AI · San Francisco, United States

External listingfull-time2 months ago

About The Role

Join Together AI, a company dedicated to building the best inference infrastructure for voice applications. As a Staff Machine Learning Engineer, you will drive the model serving layer for voice workloads, optimizing how we serve models like Whisper, Parakeet, Orpheus, and Kokoro. You will work hands-on with inference engines, profile GPU utilization, design batching strategies for streaming audio, and ensure new model architectures can go from research to production quickly. This is a foundational hire on a small, high-impact team, and you will have the opportunity to shape how Together serves voice models as the industry moves towards end-to-end speech-to-speech.

  • Own the model serving stack that powers Together's voice platform across STT, TTS, and speech-to-speech.
  • Drive best-in-class inference performance — architect and implement systems targeting leading TTFB, throughput, and GPU utilization for voice workloads.
  • Lead productionization of voice models at scale — design the serving architecture for serverless and dedicated endpoints, including batching strategies, streaming inference pipelines, and memory management tailored to real-time audio.
  • Strong foundation in speech and audio ML (ASR/TTS architectures, audio signal processing) — directly relevant experience is strongly preferred; exceptional ML engineering fundamentals with genuine curiosity about the domain is also considered
  • Sharp product intuition for developer tooling — you understand what voice application developers actually need to ship great products, and you let that shape your technical priorities, not just the other way around
  • 8+ years of ML engineering experience, with a demonstrated focus on model serving, inference optimization, or ML infrastructure at production scale — including systems you've owned from design through live traffic
  • Strong technical leadership — you operate with high autonomy, define the right problems before solving them, and raise the bar for engineering quality around you without requiring process overhead
  • Bachelor's or Master's in Computer Science, Electrical Engineering, or related field — or equivalent depth demonstrated through your work
  • Proven system design judgment — you've made architectural decisions that held up at scale and influenced how a team or platform evolved; you can articulate the tradeoffs you made and why
  • Proven ability to move fast in ambiguous environments — you've thrived on early-stage or platform teams where scope is wide, ownership is deep, and the roadmap you build is the one you execute
  • Expert-level Python and PyTorch proficiency, with a strong command of GPU optimization — CUDA kernels, memory hierarchies, profiling toolchains — and a track record of turning that knowledge into shipped latency or throughput wins
  • Familiarity with audio codec and tokenization schemes (SNAC, Encodec, DAC) is a meaningful plus at this level
  • Experience training or fine-tuning speech models at scale is a significant advantage
  • Deep, practical expertise in LLM serving engines (vLLM, SGLang, TensorRT-LLM, or equivalent) — you've modified engine internals, debugged edge cases under load, and contributed improvements back; you don't stop at the API surface

This is an external listing. JobSpring does not represent or verify the employer. Report this listing