Software Engineer (BIS (Baseten Inference Stack))
Baseten · United States
About The Role
Join Baseten, a remote-first company that is revolutionizing the way businesses deploy and operate large-scale LLM models. As a Software Engineer on the Inference Stack team, you will work across the stack, from customer-facing features to low-level infrastructure components. You will develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference, improve the reliability, scalability, and usability of our inference stack, and collaborate closely with Model Performance engineers. Enjoy unlimited PTO, full healthcare coverage, paid parental leave, and a learning and development budget.
- Développer des infrastructures et des systèmes d'orchestration pour le déploiement et la gestion de l'inférence LLM distribuée à grande échelle.
- Collaborer étroitement avec les ingénieurs en performance des modèles pour rendre les nouvelles optimisations d'inférence largement disponibles.
- Déboguer des systèmes de production complexes couvrant Kubernetes, les environnements d'exécution distribués, le réseau et les charges de travail GPU.
- Experience building and operating production systems where reliability, latency, and scale are first-class concerns
- Ability to debug complex systems across multiple layers of the stack
- Motivated and willing to learn new languages, frameworks, and systems as needed
- Strong sense of developer experience: you think about how systems are used, not just how they work
- Genuine interest in inference engineering. You don’t need to have hands on experience but are willing to learn
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, or a related field
- Strong background in distributed systems, backend infrastructure, or platform engineering
- Excellent communication and collaboration skills
- Experience contributing to open-source infrastructure or ML systems
- Experience operating GPU workloads in production
- Experience with distributed scheduling, autoscaling, or service orchestration
- Familiarity with observability tooling, CI/CD systems, or release automation
- Prior work on Dynamo, vLLM, SGLang, TensorRT-LLM, or similar inference frameworks
- Experience with Kubernetes, including concepts like operators and custom resources
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring