Skip to content
← Back to job listings

Senior Software Engineer (Machine Learning Infrastructure)

Handshake · San Francisco, United States

External listingfull-time19 days ago

About The Role

Join Handshake, a leading career marketplace, as a Senior Software Engineer in our ML Infrastructure & Platform team. In this infrastructure-heavy role, you will build and operate the shared infrastructure behind our production ML and AI systems, develop and scale our LLM platform, and optimize inference infrastructure. You will collaborate with AI, Data Science, and Product teams to productionize new models and improve the reliability, scalability, and developer experience of our ML platform. Enjoy comprehensive benefits, including health insurance, mental health support, 401k match, equity, and flexible time off.

  • Construire et exploiter l'infrastructure partagée derrière la production de l'apprentissage automatique et de l'IA, y compris les pipelines de données, les magasins de fonctionnalités, l'entraînement et le service des modèles.
  • Développer et mettre à l'échelle notre plateforme LLM, y compris les intégrations de fournisseurs, l'orchestration, l'observabilité et les contrôles pour le coût, la latence et la fiabilité.
  • Collaborer avec les équipes d'IA, de science des données et de produit pour mettre en production de nouveaux modèles et établir les meilleures pratiques pour l'infrastructure ML.
  • Strong experience with Kubernetes, Docker, Terraform, CI/CD, and operating production services
  • Experience with modern data platforms such as BigQuery, Airflow, Spark, Beam/Dataflow, or streaming pipelines
  • Strong systems design skills, sound engineering judgment, and the ability to thrive in ambiguous, fast-moving environments
  • 5+ years of production software engineering experience using Python, Go, TypeScript, or similar languages
  • Practical experience building production systems with LLMs or generative AI, including orchestration, provider APIs, observability, and performance optimization
  • Hands-on experience building ML infrastructure, including model serving, training pipelines, feature stores, embeddings, or ML observability
  • Experience building and operating cloud infrastructure on AWS, GCP, or similar platforms
  • Experience with post-training techniques such as fine-tuning, RLHF, reinforcement learning, or reward modeling
  • Experience with Vertex AI, Bigtable, Redis, or feature platform infrastructure
  • Experience designing LLM evaluation frameworks, benchmarking systems, or quality regression testing
  • Experience building agentic systems, MCP integrations, tool use, memory systems, or voice AI applications
  • Experience with Ray, Anyscale, KubeRay, Ray Serve, vLLM, Triton, PyTorch, or GPU-backed inference and training

This is an external listing. JobSpring does not represent or verify the employer. Report this listing