Skip to content
← Back to job listings

Senior Software Engineer (Machine Learning Platform)

Chime · San Francisco, United States

External listingfull-time3 months ago

About The Role

Join our team as a Machine Learning Platform Engineer, where you will design and build scalable systems that support model training, feature computation, real-time inference, and experimentation. You will work at the intersection of distributed systems, cloud infrastructure, and applied machine learning, focusing on building robust foundations that allow ML teams to move quickly while maintaining reliability, governance, and cost efficiency. Your responsibilities will include developing distributed training and batch processing systems, building and maintaining infrastructure-as-code, and improving CI/CD workflows for ML models and platform components. You will also partner closely with Data Science and ML Engineering teams to improve developer experience and contribute to platform architecture decisions and technical roadmaps.

  • Design and build scalable systems that support model training, feature computation, real-time inference, and experimentation.
  • Develop distributed training and batch processing systems using Ray, and build and maintain infrastructure-as-code using Terraform.
  • Partner closely with Data Science and ML Engineering teams to improve developer experience and contribute to platform architecture decisions.
  • Familiarity with infrastructure-as-code (e.g., Terraform, CloudFormation)
  • Strong programming skills in Python, Go, Scala, Java or similar languages
  • Strong foundation in computer science and software engineering principles
  • Deeply interested in the impact and evolution of advanced AI technologies
  • Nice-to-have
  • Solid understanding of software engineering fundamentals (testing, version control, code review, observability)
  • Experience with distributed compute frameworks such as Ray
  • Knowledge of the machine learning model development lifecycle, including data preprocessing, model training, evaluation, and deployment
  • Experience with distributed systems, cloud computing, or large-scale data processing
  • Familiarity with streaming technologies (Kafka, Kinesis, Flink, Spark Streaming, etc.)
  • Experience supporting ML lifecycle workflows (training, evaluation, deployment, monitoring)
  • 5+ years of experience in ML infrastructure, platform engineering, or production ML systems
  • Knowledge of ML experimentation platforms and model governance practices
  • Hands-on experience with CI/CD pipelines, DevOps practices, and infrastructure as code
  • Experience building or operating a feature store
  • Knowledge of cloud platforms such as AWS and distributed computing frameworks such as Spark and Ray
  • Experience with containerization technologies such as Docker and Kubernetes, and orchestration systems
  • Experience with GPU programming(CUDA) and GPU costs/optimization
  • Experience with real-time ML systems or model serving

This is an external listing. JobSpring does not represent or verify the employer. Report this listing