← Back to job listings
AN
Staff+ Software Engineer (Capacity Engineering)
Anthropic · New York, United States
About The Role
Join Anthropic, a leading AI safety and research company. As a Staff+ Software Engineer in Capacity Engineering, you will build production systems that optimize infrastructure utilization and efficiency. You will work closely with various teams and have a direct impact on one of Anthropic's largest areas of spending. This role requires expertise in data engineering, systems engineering, and observability, and offers a high level of autonomy and ambiguity.
- Construire des systèmes de production qui alimentent le travail de l'équipe d'ingénierie de capacité, y compris des pipelines de données qui ingèrent et normalisent la télémétrie.
- Collaborer étroitement avec les équipes d'ingénierie de recherche, d'infrastructure, d'inférence et de finance pour maximiser l'utilisation des ressources.
- Développer des outils de planification et d'allocation, des programmes d'efficacité et des systèmes d'attribution et de prévision.
- A strong track record building and operating production systems. This is a hands-on engineering role with a devops flavor
- Deep experience with at least one major cloud provider (Amazon Web Services, Google Cloud, or Microsoft Azure) and its operations
- Experience with observability tooling stack, including Prometheus, PromQL, and Grafana, including writing recording rules and building monitoring that engineering teams rely on
- Python and SQL at production quality. Most pipeline code is Python; the presentation layer is BigQuery SQL, including table-valued functions and views. Both need to be idiomatic, well-tested, and maintainable
- Ability to gather your own requirements and work across organizational boundaries in an ambiguous environment with limited direction
- Experience with capacity planning, resource management, or cost attribution systems at a hyperscaler or in a large-scale machine learning environment. Time spent in product engineering and developer experience absolutely counts here
- Scheduling and packing efficiency experience, or profiling-driven optimization of large distributed workloads
- Accelerator infrastructure familiarity. GPU metrics (DCGM), TPU utilization, Trainium power and utilization metrics, or experience with machine learning training and inference systems at the hardware level
- Multi-cloud data ingestion experience, especially normalizing billing exports, reservation APIs, on-demand capacity reservations, commitments, and vendor telemetry from providers with different billing arrangements
- Total cost of ownership and forecasting experience, including decomposing whether infrastructure growth is causal or correlated with business drivers
- Experience building internal data products with self-service access, schema contracts, API serving, documentation, and discoverability. Not just pipelines, but thinking about how the data gets consumed
- Storage efficiency, retention, and lifecycle program experience at exabyte scale
- Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
- Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
- Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices
- We encourage you to apply even if you do not believe you meet every single qualification
- Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring