Skip to content
← Back to job listings

Senior Machine Learning Engineer (Inference Platform)

Wizard · United States

External listingfull-time2 months ago

About The Role

Join Wizard AI, a leading AI Shopping Agent company, as a Senior Machine Learning Engineer. In this role, you will own the end-to-end lifecycle of production ML serving systems, including model packaging, deployment, monitoring, optimization, and scaling. You will work closely with ML Engineers, Data teams, Product, and DevOps to ensure seamless model transitions from experimentation to high-performance production systems. Your responsibilities will include evolving the multi-engine inference platform, building and improving production ML pipelines, defining and implementing model versioning and lifecycle management strategies, and optimizing inference performance. You will also build observability, monitoring, alerting, and operational tooling for production inference systems.

  • Posséder l'ensemble du cycle de vie de la production des systèmes de service ML, de l'emballage et du déploiement des modèles à la surveillance, à l'optimisation et à l'évolutivité.
  • Être responsable de l'infrastructure d'inférence qui alimente un agent de shopping conversationnel en direct, en faisant fonctionner plusieurs moteurs de service spécialisés sous une charge de production réelle.
  • Définir et mettre en œuvre la version des modèles, le déploiement, le retour en arrière et les stratégies de gestion du cycle de vie qui garantissent la reproductibilité et la fiabilité opérationnelle.
  • Bachelor's or Master's degree in Computer Science, Data Science, Engineering, or a related field, or equivalent practical experience
  • Demonstrated ability to balance latency, throughput, reliability, and infrastructure cost while operating production-scale ML systems
  • Experience serving heterogeneous workloads, including LLMs, embedding models, and extraction models, each with distinct latency, throughput, and scaling requirements
  • 5–8+ years of experience in Software Engineering, ML Engineering, Platform Engineering, or Infrastructure Engineering, with direct ownership of production ML serving systems
  • Strong Python skills and software engineering fundamentals, combined with deep systems and infrastructure knowledge
  • Strong grasp of inference performance — continuous batching, KV-cache and GPU-memory behavior, quantization, and CPU-versus-GPU bottlenecks — with the instinct to profile before tuning
  • Experience with cloud platforms such as AWS, GCP, or Azure, and familiarity with ML lifecycle tooling, experimentation platforms, and model registries
  • Experience in high-growth startup environments and comfort operating in fast-moving, evolving technical landscapes
  • Hands-on experience running an LLM serving engine (vLLM, TGI, TensorRT-LLM, or SGLang) in production under real load — not just managed or hosted endpoints

This is an external listing. JobSpring does not represent or verify the employer. Report this listing