Skip to content
← Back to job listings

Software Engineer (ML Infrastructure)

Nuro · Mountain View, CA, United States

External listingfull-timeabout 1 month ago

About The Role

Join Nuro, a leading robotics company, as a Software Engineer specializing in ML Infrastructure. In this role, you will build and evolve the core platform that provides researchers and engineers with seamless access to compute and data resources. You will be responsible for executing the technical strategy for automated resource provisioning, high-performance workload scheduling, and efficient feature management. This position offers a range of benefits, including a free Caltrain pass, company stock options, work from home opportunities, and health insurance.

  • Construire et faire évoluer la plateforme centrale qui fournit aux chercheurs et aux ingénieurs un accès fluide aux ressources de calcul et de données.
  • Exécuter la stratégie technique pour l'approvisionnement automatisé des ressources, la planification des charges de travail haute performance et la gestion efficace des fonctionnalités.
  • Contribuer à une plateforme ML unifiée qui abstrait l'infrastructure cloud complexe pour les utilisateurs finaux.
  • Systems Design: A strong understanding of distributed systems, networking, and storage bottlenecks in the context of high-performance computing
  • Resource Provisioning: Deep familiarity with modern Infrastructure-as-Code and provisioning tools such as Terraform, Pulumi, or Crossplane
  • Distributed Data Processing: Proficiency in at least one distributed processing framework, such as Apache Spark or Apache Beam, for large-scale data extraction and transformation
  • Experience: 3+ years of professional experience in ML Infrastructure, Backend Platform Engineering, or Distributed Systems
  • Feature Management: Experience implementing or maintaining feature stores and caching layers (e.g., Feast, Hopsworks, or Redis-based custom caching)
  • Workload Scheduling: Hands-on experience building or managing large-scale orchestrators for compute-heavy workloads (e.g., Kubernetes, KubeRay, Ray, Slurm, or Volcano)
  • Active contributor to open-source projects in the MLOps or Cloud-Native ecosystem (e.g., CNCF, Ray, or Kubeflow communities)
  • Experience with high-performance storage systems (e.g., Lustre, Ceph, or specialized NVMe caching) for ML data loading
  • Knowledge of cost-optimization strategies for large-scale GPU clusters in public clouds (AWS, GCP, or Azure)

This is an external listing. JobSpring does not represent or verify the employer. Report this listing