Skip to content
← Back to job listings

Engineering Manager (Forward Deployed Engineering, LLM)

Baseten · United States

External listingfull-time3 days ago

About The Role

Join Baseten as an Engineering Manager, where you will lead and mentor a team of Forward Deployed Engineers focused on building, scaling, and optimizing LLM inference workloads. You will guide your team through the processes of designing, deploying, and managing high-performance, low-latency AI applications on Baseten’s platform. This role requires a mix of hands-on technical ownership and managerial leadership, as well as collaboration with product, infrastructure, and other customer engineering teams.

  • Lead and mentor a team of Forward Deployed Engineers in building, scaling, and optimizing LLM inference workloads for customers.
  • Guide the team through the processes of designing, deploying, and managing high performance, low latency AI applications on Baseten’s platform.
  • Collaborate with leadership to align team priorities with company and customer goals, balancing short-term delivery and long-term technical initiatives.
  • Strong programming skills in Python, with production experience in building or optimizing ML inference systems
  • Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, or related field
  • Proven experience with LLMs, inference optimization, or serving frameworks (e.g., vLLM, TensorRT, Triton, Hugging Face, Ray Serve)
  • Familiarity with observability, profiling, and cost/performance tradeoffs in production ML systems
  • 4+ years of professional software engineering experience, including 1+ year in a leadership or mentorship capacity
  • Excellent communication and collaboration skills—able to lead cross-functional efforts and drive outcomes in ambiguous, fast-paced environments
  • If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you
  • Experience leading customer-facing engineering teams or working directly with enterprise partners
  • Deep understanding of GPU infrastructure, distributed inference, or model compression techniques

This is an external listing. JobSpring does not represent or verify the employer. Report this listing