Skip to content
← Back to job listings

Senior Backend Engineer (Inference Platform)

Together AI · San Francisco, United States

External listingfull-time22 days ago

About The Role

Join our team as a Senior Backend Engineer, where you'll work hands-on with cutting-edge GPU technology and collaborate directly with research teams to bring frontier models into production. You'll be responsible for optimizing global and local request routing, developing auto-scaling systems, and designing systems for multi-tenant traffic shaping. This role requires expert-level programming skills, a strong understanding of low-level OS concepts, and experience building large-scale, fault-tolerant distributed systems. Enjoy competitive health insurance, flexible time off, and a collaborative work environment.

  • Collaborate with research teams to bring frontier models into production and contribute to open source projects.
  • Build and optimize global and local request routing, ensuring low-latency load balancing across data centers.
  • Develop auto-scaling systems to dynamically allocate resources and meet strict SLOs across dozens of data centers.
  • If you get a thrill from optimizing latency down to the last millisecond, this is your playground
  • Experience with Kubernetes or container orchestration is a strong plus
  • Expert-level programming in one or more of: Rust, Go, Python, or TypeScript
  • Excellent understanding of low-level OS concepts: multi-threading, memory management, networking, and storage performance
  • Experience working with the open source ecosystem around inference is highly valuable; familiarity with SGLang, vLLM, or NVIDIA Dynamo will be especially handy
  • Knowledge of modern LLMs and generative models and how they are served in production is a plus
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience
  • Familiarity with GPU software stacks (CUDA, Triton, NCCL) and HPC technologies (InfiniBand, NVLink, MPI) is a plus
  • 5+ years of demonstrated experience building large-scale, fault-tolerant, distributed systems and API microservices
  • Strong background in designing, analyzing, and improving efficiency, scalability, and stability of complex systems

This is an external listing. JobSpring does not represent or verify the employer. Report this listing