Skip to content
← Back to job listings

Engineering Manager (Model Inference)

Abridge · San Francisco, United States

External listingfull-time3 months ago

About The Role

Join Abridge, a fast-growing startup revolutionizing the practice of medicine with generative AI-powered products. As an Engineering Manager, you will lead and grow the Model Inference team, owning the technical direction of our inference systems and ensuring peak efficiency and reliability. You will work closely with various teams, including ML Research, GenAI Platform, Data, and Product, to plan and execute projects that directly impact clinicians and patients. Enjoy a remote work environment, equity for all new employees, unlimited PTO, and a generous equipment budget for your home office setup.

  • Lead and grow a high-performing team of AI inference engineers focused on building and scaling infrastructure for Abridge’s products and APIs.
  • Own the technical direction of our inference systems—making key decisions around batching, throughput, latency, and GPU utilization.
  • Architect and scale inference infrastructure for reliability, efficiency, and observability; lead incident response.
  • Experience with inference optimizations (eg. batching, quantization, kernel fusion, FlashAttention)
  • Strong understanding of LLM architecture (eg. Multi-Head Attention, Multi/Grouped-Query Attention, and common transformer components)
  • Experience with parallelism strategies: tensor parallelism, pipeline parallelism, expert parallelism
  • Strong technical communication and cross-functional collaboration skills
  • Has thrived in a fast-growing startup and knows how to operate with urgency and focus
  • Experience deploying reliable, distributed, real-time systems at scale
  • 5+ years of engineering experience with 1+ years in a technical leadership or management role
  • Familiarity with GPU characteristics, roofline models, and performance analysis
  • Deep, hands-on experience with ML systems and inference frameworks (e.g., PyTorch, TensorRT, vLLM, TensorFlow)
  • Comfortable giving constructive feedback on technical designs and code reviews
  • Skilled at hiring and mentorship, with a demonstrated track record of helping engineers grow their skills and careers
  • Background in training infrastructure and RL workloads
  • Skilled in building secure, compliant systems on major cloud platforms (GCP preferred, AWS experience welcome)
  • Experience with Kubernetes and container orchestration at scale
  • Published work or contributions to inference optimization research

This is an external listing. JobSpring does not represent or verify the employer. Report this listing