Skip to content
← Back to job listings

Machine Learning Infrastructure Engineer (Model Inference)

Abridge · San Francisco, United States

External listingfull-time25 days ago

About The Role

Join Abridge as a Machine Learning Infrastructure Engineer, where you'll be instrumental in building and optimizing the core inference infrastructure for our AI-driven solutions. You'll collaborate with our Infrastructure and Research teams to enhance scalability, efficiency, and performance. This remote position offers unlimited PTO, 16 weeks of paid parental leave, equity for all new employees, and a generous equipment budget for your home office setup.

  • Conception, déploiement et maintenance de clusters Kubernetes évolutifs pour l'inférence et l'entraînement des modèles d'IA.
  • Développement, optimisation et maintenance de l'infrastructure de service des modèles d'apprentissage automatique, garantissant des performances élevées et une faible latence.
  • Collaboration avec les équipes de ML et de produit pour faire évoluer l'infrastructure backend pour les produits alimentés par l'IA, en se concentrant sur le déploiement des modèles, l'optimisation du débit et l'efficacité des calculs.
  • Deep understanding of container orchestration and distributed systems architecture
  • Expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management
  • Excellent communication skills, with the ability to interface between research and product engineering
  • Experience developing APIs and managing distributed systems for both batch and real-time workloads
  • 2+ years of experience in building and deploying machine learning models in production environments
  • Expertise with model serving frameworks such as NVIDIA Triton Server, VLLM, TRT-LLM and so on
  • Knowledge of infrastructure as code (Terraform, Ansible) and GitOps practices
  • Familiarity with GPU cluster management and CUDA optimization
  • Expertise with ML toolchains such as PyTorch, Tensorflow or distributed training and inference libraries
  • Experience with container registries, image optimization, and multi-stage builds for ML workloads
  • Experience orchestrating across ASR models or LLM models for building various GenAI applications

This is an external listing. JobSpring does not represent or verify the employer. Report this listing