← Back to job listings
KA
Senior Machine Learning Engineer
Kargo · New York, United States
About The Role
Join Kargo, a leading mobile advertising platform, as a Senior Machine Learning Engineer. In this role, you will lead the design and production deployment of multimodal ML models that quantify creative quality and predict ad performance. You will be the technical anchor for the Creative Sciences Platform, translating research in LLMs, VLMs, and multimodal learning into scalable, reliable systems. Your success will be measured by the improved predictive accuracy of the creative scoring model, the establishment of production-grade MLOps, and the operationalization of model reliability.
- Lead the design and production deployment of multimodal ML models that quantify creative quality and predict ad performance.
- Establish end-to-end pipelines for model training, fine-tuning, deployment, and monitoring, ensuring rapid iteration from notebook to production.
- Build and operate the APIs, embedding services, and model endpoints that allow other platforms to consume scoring in real time.
- Expert in Python and PyTorch (or TensorFlow), plus distributed training frameworks (Ray, PyTorch Lightning, Horovod)
- Strong SQL, data pipeline, and feature store design for scalable experimentation
- 5+ years in ML engineering or MLOps, with shipped production systems involving LLMs, VLMs, or multimodal architectures
- Cloud-native ML deployment on AWS (SageMaker), GCP (Vertex AI), or Azure ML, with infrastructure-as-code (Terraform, Helm)
- Production fluency with Docker, Kubernetes, and CI/CD patterns for ML
- Hands-on with MLOps tooling: MLflow, Weights & Biases, Kubeflow, Argo, or Airflow for orchestration, experiment tracking, and automated retraining
- Experience with vector databases, embedding pipelines, and real-time retrieval systems
- Background in creative scoring, aesthetic modeling, or ad performance prediction
- Translates papers and prototypes into systems that survive production traffic, monitoring, and on-call
- Knows when a model is good enough to ship vs. when it needs another iteration — doesn't over-engineer or under-validate
- Designs for the second and third version of the model, not just the first — pipelines, abstractions, and infra that compound over time
- Optimizes the full stack: training cost, inference latency, and developer iteration speed, not just model accuracy
- Explains multimodal modeling tradeoffs to Product and Creative stakeholders in terms of business impact, not architecture diagrams
- Partners with Data Science and Platform Engineering as co-owners, not handoff points
- Treats drift, latency regressions, and silent failures as personal — instruments systems so problems are caught early and root-caused fast
- Documents architecture, decisions, and runbooks so the platform outlives any single contributor
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring