Product Manager, Compute Platform
San Francisco, CA | New York City, NY | Seattle, WA
Posted 16d ago
Job Location
San Francisco, CA | New York City, NY | Seattle, WA
Tech Stack
Remote Work Policy
On-site
Categories
AI Infrastructure Engineer
About the job
Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. As a Product Manager focused on Compute Platform, you will partner with Infrastructure, Compute Operations, Engineering, Finance & Strategy, and Research teams. Your role will be to build the scheduling, orchestration, and capacity management systems that power Anthropic’s compute infrastructure, which is essential for all model training, evaluation, and inference workloads.
Responsibilities
- Partner with Infrastructure to build systems for scheduling, prioritizing, and allocating jobs across GPU and accelerator clusters.
- Define the semantic layer for job scheduling, establish resource guarantees, and manage trade-offs for peak infrastructure capacity.
- Drive the evolution of the compute platform to support diverse workloads like large-scale training, fine-tuning, real-time inference, and batch evaluation.
- Define and own the strategy and roadmap for job scheduling primitives, capacity allocation policies, preemption and fairness frameworks, quota management, and observability tooling.
- Understand the needs of internal customers across Research, Infrastructure, Product, and Finance.
- Define and iterate on the semantic layer for job scheduling, including abstractions, priority tiers, resource classes, and preemption policies.
- Partner with engineering leads to design scheduling capabilities that maximize cluster utilization while honoring resource guarantees.
- Drive product strategy and roadmap for compute capacity management, including quota systems, fairness policies, bin-packing optimizations, and gang-scheduling.
- Own the trade-off framework between utilization efficiency, job latency, cost, and reliability.
- Collaborate with the Capacity Strategy & Operations team on capacity planning models, demand forecasting, and cost-to-serve analytics.
- Build and champion observability tools and dashboards for real-time visibility into cluster health, queue depth, scheduling efficiency, and resource waste.
Requirements
- 7+ years of product management experience with deep exposure to compute infrastructure, distributed systems, or scheduling/orchestration platforms.
- Experience taking technical infrastructure products from infancy to scale.
- Track record of building platform products that balance the needs of multiple users and stakeholders.
- Ability to internalize complex technical systems and translate that understanding into a comprehensive product vision.
- Credible discussing scheduling algorithms with engineers, capacity economics with finance, and infrastructure strategy with leadership.
- Strong instinct for connecting technical decisions to business outcomes.
- Scrappy and resourceful, able to get things done in a fast-moving environment.