← Back to job listings
LA
Staff Product Manager (Observability)
Lambda · United States
About The Role
Join Lambda, a leading AI infrastructure company, as a Staff Product Manager for Observability. In this role, you will own the observability product across Lambda's cloud, partnering with various teams to deliver a coherent experience for customers and operators. You will define the product strategy, turn telemetry into product features, and work closely with customers to understand their needs. This position offers a competitive salary, benefits, and the opportunity to make a significant impact in the AI industry.
- Define and drive the product strategy for observability across Lambda's cloud, from single On-Demand GPU Instances to 1-Click Clusters at 64 to 1,024+ GPU scale.
- Decide which signals customers see, including GPU and cluster health, utilization, job-level telemetry, and InfiniBand fabric metrics, and how they see them in the console and the Lambda Cloud API.
- Partner with SRE and fleet engineering to define the internal telemetry that powers fleet operations and reliability, so one data foundation serves both customers and operators.
- If you love turning raw telemetry into products customers rely on, and you want your work to be the reason a research team trusts their 1,024-GPU training run, we'd love to hear from you
- We value diverse backgrounds, experiences, and skills, and we are excited to hear from candidates who can bring unique perspectives to our team
- If you do not exactly meet this description but believe you may be a good fit, please still apply and help us understand your readiness for this role
- Are energized by ambiguity, and can be the first product manager to own a domain and give it shape
- Can write crisply, so a one-page document from you is enough to align a room
- Able to define iterative plans that move an organization from the current state towards the desired outcome
- Have taken products from idea through launch, measurement, and iteration, and can speak to what you learned from the metrics after shipping
- Are comfortable in deeply technical conversations about distributed systems, and can hold your own with engineers on topics like metrics pipelines, hardware health, and failure modes
- Have 7+ years of product management experience, including 3+ years on technical infrastructure, platform, or developer-facing products
- Turn data and customer signal into a clear decision about what to build next, and can show examples where your insight changed a roadmap
- Win over engineers, designers, executives, and customers without formal authority, and have shipped products that required cross-team adoption
- Have shipped monitoring or observability products, such as those in the Datadog, Grafana, or Prometheus class
- Have worked with HPC (high performance computing) or distributed systems telemetry, including interconnect and fabric-level metrics
- Have hands-on exposure to distributed training stacks such as PyTorch with NCCL, or to GPU fleet tooling such as DCGM (Data Center GPU Manager)
- Have built developer tools or API-first products
- Have worked in a usage-based cloud infrastructure business
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring