← Back to job listings
TA
Multimodal AI Model Optimization Research Engineer
Tavus · United States
About The Role
Join our core AI team as a Multimodal AI Model Optimization Research Engineer. In this role, you will focus on optimizing cutting-edge research models for speed, efficiency, and production readiness. You will own the optimization lifecycle for key models, define metrics, run experiments, and benchmark trade-offs across latency, cost, and quality. You will collaborate closely with researchers and engineers to turn new ideas into deployable systems. This position offers comprehensive benefits, unlimited paid time off, flexible hours, and a remote-friendly work environment.
- Prendre des modèles de recherche de pointe et les rendre rapides, efficaces et prêts pour la production en utilisant la sparsification, la distillation et la quantification.
- Posséder le cycle de vie de l'optimisation pour des modèles clés : définir des métriques, exécuter des expériences et évaluer les compromis entre la latence, le coût et la qualité.
- Collaborer étroitement avec les chercheurs et les ingénieurs pour transformer de nouvelles idées en systèmes déployables.
- Our ideal partner-in-crime thrives in startup environments, is comfortable prioritizing independently, and is willing to take calculated risks
- We’re moving fast and looking for people who can help pave the path
- Hands-on experience with model optimization and compression, including knowledge distillation, pruning/sparsification, quantization, and mixed precision
- Ability to read ML papers, reproduce results, and adapt ideas
- Clear communication and collaboration skills
- Experience working with large models and datasets in cloud environments
- Strong experience in deep learning using PyTorch
- Strong Python coding skills and reliable research engineering practices
- Strong understanding of inference performance and GPU/accelerator fundamentals
- Understanding of efficient architectures such as low-rank adapters
- Experience with real-time or streaming systems (low-latency APIs, WebRTC, streaming TTS/video)
- Optimization of diffusion models, video/audio generative models, or large language models
- Familiarity with TensorRT, ONNX Runtime, TVM, Triton, or XLA
- Experience writing custom Triton/CUDA kernels or low-level performance tuning
- Experience with experiment tracking, benchmarking, and profiling at scale
- Prior experience in research engineering or applied science roles
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring