← Back to job listings
BA
Engineering Manager (Forward Deployed Engineering, LLM)
Baseten · United States
About The Role
Join Baseten as an Engineering Manager, where you will lead and mentor a team of Forward Deployed Engineers focused on building, scaling, and optimizing LLM inference workloads. You will guide your team through the processes of designing, deploying, and managing high-performance, low-latency AI applications on Baseten’s platform. This role requires a mix of hands-on technical ownership and managerial leadership, as well as collaboration with product, infrastructure, and other customer engineering teams.
- Lead and mentor a team of Forward Deployed Engineers in building, scaling, and optimizing LLM inference workloads for customers.
- Guide the team through the processes of designing, deploying, and managing high performance, low latency AI applications on Baseten’s platform.
- Collaborate with leadership to align team priorities with company and customer goals, balancing short-term delivery and long-term technical initiatives.
- Strong programming skills in Python, with production experience in building or optimizing ML inference systems
- Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, or related field
- Proven experience with LLMs, inference optimization, or serving frameworks (e.g., vLLM, TensorRT, Triton, Hugging Face, Ray Serve)
- Familiarity with observability, profiling, and cost/performance tradeoffs in production ML systems
- 4+ years of professional software engineering experience, including 1+ year in a leadership or mentorship capacity
- Excellent communication and collaboration skills—able to lead cross-functional efforts and drive outcomes in ambiguous, fast-paced environments
- If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you
- Experience leading customer-facing engineering teams or working directly with enterprise partners
- Deep understanding of GPU infrastructure, distributed inference, or model compression techniques
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring