Skip to content
← Back to job listings

VP of Engineering

Hyperbolic · San Francisco, United States

External listingfull-timeabout 2 months ago

About The Role

Join our team as the Vice President of Engineering, where you will lead the design and evolution of our AI cloud platform. This hands-on executive leadership role requires deep engagement in architecture, design, and engineering execution. You will define the architecture for GPU orchestration, compute scheduling, networking, storage, and distributed systems, and make critical decisions regarding cloud infrastructure and platform scalability. You will also build and scale large GPU clusters, drive platform reliability and performance, and establish best practices for Kubernetes, observability, CI/CD, security, and operational excellence.

  • Lead the design and evolution of the AI cloud platform, defining the architecture for GPU orchestration, compute scheduling, networking, storage, and distributed systems.
  • Build and scale large GPU clusters supporting customer workloads, driving platform reliability and performance for AI training and inference workloads.
  • Recruit and develop world-class Infrastructure, Platform, and SRE teams, building a high-performance engineering culture focused on ownership and execution.
  • Deep expertise in observability, monitoring, and reliability engineering
  • Previous experience building or operating a cloud platform at scale
  • Experience leading infrastructure organizations while remaining hands-on technically
  • Expert-level Kubernetes knowledge
  • 12+ years building and operating large-scale infrastructure systems
  • Experience building highly available production systems
  • Experience building GPU infrastructure or AI/ML compute platforms
  • Experience designing and operating multi-region cloud infrastructure
  • Strong understanding of Linux, networking, distributed systems, and storage architecture
  • Experience with Infrastructure-as-Code and automation frameworks
  • Proven track record scaling infrastructure in high-growth startup environments
  • Background supporting AI training and inference platforms
  • Experience managing thousands of GPUs in production environments
  • Experience with GPU scheduling, Slurm, Kubernetes GPU operators, Ray, or distributed training systems

This is an external listing. JobSpring does not represent or verify the employer. Report this listing