Product Manager (Compute Platform)
Anthropic · New York, United States
About The Role
Join Anthropic as a Product Manager focused on Compute Platform. In this role, you will partner with various teams to build the scheduling, orchestration, and capacity management systems that power Anthropic's compute infrastructure. Your work will directly impact cluster utilization, cost efficiency, and researcher velocity. You will define and own the strategy and roadmap for job scheduling primitives, capacity allocation policies, and observability tooling. Additionally, you will collaborate with internal customers across Research, Infrastructure, Product, and Finance to understand their needs and drive product strategy.
- Collaborer avec les équipes d'infrastructure pour construire des systèmes de planification, d'orchestration et de gestion de la capacité qui optimisent l'utilisation des clusters de calcul.
- Définir et posséder la stratégie et la feuille de route pour les primitives de planification des tâches, les politiques d'allocation de capacité, et les outils d'observabilité.
- Comprendre les besoins des clients internes et itérer sur la couche sémantique de la planification des tâches, en établissant des garanties de ressources et en prenant des décisions de priorisation transparentes.
- 7+ years of product management experience, with deep exposure to compute infrastructure, distributed systems, or scheduling/orchestration platforms
- Track record of building platform products that balance the needs of multiple users and stakeholders—you’re comfortable making prioritization trade-offs between utilization, latency, cost, and fairness, and communicating them clearly
- Strong instinct for connecting technical decisions to business outcomes: every percentage point of cluster utilization has measurable impact
- Experience taking technical infrastructure products from infancy to scale—you’ve built something from the ground up and grown it to serve demanding internal or external customers
- Scrappy and resourceful—you do what it takes to get things done in a fast-moving environment
- Ability to internalize complex technical systems (job schedulers, cluster managers, resource orchestrators) and translate that understanding into a comprehensive product vision
- Fluent across functions—you’re equally credible discussing scheduling algorithms with engineers, capacity economics with finance, and infrastructure strategy with leadership
- Scaled through hypergrowth in compute-intensive environments (AI/ML, HPC, large-scale cloud infrastructure)
- Built or scaled job scheduling, resource orchestration, or workload management systems for large-scale compute clusters (e.g., Kubernetes, Slurm, Borg, YARN, or custom schedulers)
- Experience with observability and efficiency tooling for distributed infrastructure—building dashboards, automation, and governance workflows that drive utilization and cost accountability
- Capacity planning experience across cloud and on-premises infrastructure, including cost modeling, demand forecasting, and vendor management for compute procurement
- Not all strong candidates will meet every single qualification as listed
- We encourage you to apply even if you do not believe you meet every single qualification
- Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work
- Experience defining and enforcing SLAs and resource guarantees for compute workloads—including mechanisms to validate job prerequisites (data readiness, checkpoint availability, hardware compatibility) before scheduling to avoid wasted resources
- Deep familiarity with GPU/accelerator scheduling challenges, including gang-scheduling, topology-aware placement, preemption, and hardware affinity constraints
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring