Skip to content
← Back to job listings

Head of Platform Product Reliability

Etched · San Jose, United States

External listingfull-time21 days ago

About The Role

Join Etched, a leading company in AI infrastructure. As the Head of Platform Product Reliability, you will lead reliability engineering across our server, rack, and datacenter platform products. You will define and own the end-to-end reliability strategy, establish reliability requirements and validation methodologies, and build a high-performing product reliability engineering organization. This role requires a deep understanding of system-level failure mechanisms, hands-on experience with reliability analysis methodologies, and a track record of driving cross-functional root-cause investigations.

  • Definir y ser responsable de la estrategia de confiabilidad de los productos de plataforma, desde los requisitos de diseño hasta el despliegue en el campo.
  • Establecer y institucionalizar procesos de ingeniería de confiabilidad que abarquen todo el ciclo de vida del producto, incluyendo pruebas de vida aceleradas y pruebas de estrés aceleradas.
  • Liderar investigaciones de causa raíz para fallas de confiabilidad y desarrollar modelos de confiabilidad del sistema, incluyendo proyecciones de MTBF y análisis de tasas de FIT.
  • Deep understanding of system-level failure mechanisms — including thermal, power delivery, mechanical, and connector/interconnect failure modes — and how design decisions affect long-term field reliability
  • Networking platforms, storage systems, or rack-scale infrastructure
  • Hands-on experience with FMEA, Weibull analysis, HALT/HASS, qualification planning, failure analysis methodologies, and reliability statistics and modeling
  • BS, MS, or PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering, or a related technical field
  • Excellent communication skills and the credibility to influence design decisions with engineering leads, program managers, and executive stakeholders
  • 10+ years of reliability engineering experience in hardware-centric organizations, with meaningful time spent on complex systems rather than component-level work
  • A track record of driving cross-functional root-cause investigations in fast-moving hardware organizations where schedule pressure is real and accountability is high
  • Strong technical judgment — capable of making defensible tradeoffs between reliability targets, cost, schedule, and performance without losing sight of customer expectations
  • Experience leading reliability programs for one or more of:
  • AI accelerator or GPU-class compute systems
  • Hyperscale or cloud server infrastructure
  • Experience with liquid-cooled systems, high-density power delivery, or thermal management for high-power AI infrastructure
  • Direct experience supporting hyperscale or cloud datacenter deployments at scale, including customer-facing reliability commitments and SLA management
  • Demonstrated experience building a reliability organization from early-stage — establishing processes, tooling, and team norms in environments without established infrastructure
  • Familiarity with fleet telemetry systems, large-scale field reliability analytics, and data-driven approaches to proactive reliability management
  • Experience working closely with ODM or JDM partners in Taiwan or broader Asia, including NPI support and on-site qualification engagement
  • Background in high-speed digital systems, GPU compute platforms, or accelerator-based architectures — with an understanding of how these affect system-level reliability behavior

This is an external listing. JobSpring does not represent or verify the employer. Report this listing