Senior Machine Learning Engineer (DevOps/SRE)
Roku · San Jose, United States
About The Role
Join Roku's Advertising Performance team as a Senior Machine Learning Engineer (DevOps/SRE). In this role, you will support and scale our Machine Learning infrastructure, streamline the end-to-end ML lifecycle, and lead the design and operation of scalable cloud infrastructure for ML workloads. You will also define and enforce observability standards for ML systems and participate in on-call rotation for critical ML training and serving infrastructure. The ideal candidate has a strong background in DevOps/SRE practices, cloud infrastructure management, and MLOps tooling, with a passion for building platforms that accelerate ML experimentation and deployment at internet scale.
- Lead the design and operation of scalable, production-grade cloud infrastructure for ML workloads across AWS and GCP, including GPU/TPU-based training and inference environments.
- Architect and improve CI/CD systems for ML models and platform services to enable fast, reliable, and safe production releases.
- Define and enforce observability standards for ML systems, including model performance monitoring, drift detection, capacity planning, and pipeline health metrics.
- The ideal candidate has a strong background in DevOps/SRE practices, cloud infrastructure management, and MLOps tooling — with a passion for building platforms that accelerate ML experimentation and deployment at internet scale
- Experience in the Advertising domain is a plus
- Hands-on experience with data and orchestration technologies such as Apache Spark, Apache Flink, Apache Airflow, and Kafka
- Experience with observability platforms such as Prometheus, Grafana, and Datadog
- Expertise with NoSQL or low-latency data stores such as Aerospike or similar technologies
- Excellent communication and cross-functional collaboration skills
- Deep experience with Kubernetes and container orchestration on GCP (GKE) and/or AWS (EKS)
- BS or MS in Computer Science, Engineering, or a related quantitative field
- 8+ years of experience in DevOps, SRE, or ML infrastructure, including 4+ years supporting large-scale ML or AI systems
- Strong programming skills in Python, and/or Scala, or Java for platform automation and tooling
- Experience building and maintaining CI/CD systems using tools such as Jenkins or GitLab Runner
- Familiarity with feature engineering platforms such as Chronon and model lifecycle tools such as MLflow
- Strong infrastructure-as-code experience with Terraform or similar tooling
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring