Skip to content
← Back to job listings

Engineering Manager (Infrastructure Engineering)

CoreWeave · Sunnyvale, United States

External listingfull-time7 days ago

About The Role

Join CoreWeave, a rapidly growing AI cloud provider, as an Engineering Manager in Infrastructure Engineering. In this role, you will lead the MetalDev RAS team, responsible for the reliability, availability, and serviceability of CoreWeave's bare-metal infrastructure. You will balance hands-on technical leadership with people leadership, driving operational excellence and partnering with cross-functional teams.

  • Construire, encadrer et développer une équipe d'ingénieurs en infrastructure et en fiabilité des sites, en favorisant une culture de responsabilité et de fiabilité.
  • Assurer la fiabilité, la disponibilité et la serviceabilité des services MetalDev RedFish en production, en définissant et en pilotant les KPI, SLA et SLO.
  • Collaborer avec les équipes d'ingénierie de l'organisation sur la fiabilité de la plateforme, les améliorations de la résilience et la récupération après sinistre.
  • Solid grounding in incident management practices and frameworks (e.g., ITIL, SRE best practices), including on-call, RCA, and PIR processes
  • Strong understanding of cloud platforms and K8S, and of cloud and bare-metal infrastructure
  • 3+ years of engineering management experience leading and growing teams of software / infrastructure / SRE engineers (first-line management), on top of a strong individual-contributor background
  • 7+ years of combined experience in cloud operations, site reliability engineering (SRE), infrastructure, or related technical roles
  • Experience with observability tooling such as Prometheus and Grafana
  • Experience supporting production services on an on-call rotation, and a demonstrated ability to build a healthy on-call and operational culture within a team
  • Excellent communication, documentation, and stakeholder-management skills, with strong analytical and problem-solving abilities
  • Track record of defining and driving reliability metrics (SLAs / SLOs / KPIs) and operational improvements at scale
  • Experience leading teams that develop software in Go or (comparable systems languages) with the ability to participate in architecture and code reviews
  • Familiarity with Redfish, BMC, or server lifecycle management technologies
  • Experience managing bare-metal, hardware, or hardware-adjacent infrastructure teams
  • Experience scaling teams, tooling, and processes in a high-growth environment
  • Or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency
  • Must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. § 1157, or (iv) asylee under 8 U.S.C. § 1158, (B) eligible to access the export controlled information without a required export authorization

This is an external listing. JobSpring does not represent or verify the employer. Report this listing