Technical Program Manager, Infrastructure
San Francisco, CA | New York City, NY | Seattle, WA
Posted 16d ago
About the job
Anthropic's Infrastructure organization is the engine that powers our mission, building and operating massive clusters for training frontier models, production infrastructure serving millions of users reliably, and developer platforms. As a Technical Program Manager for Infrastructure, you will coordinate complex programs with broad organizational impact, solving novel scaling challenges while maintaining security and reliability. This role is ideal for someone who thrives in ambiguity, makes others more effective, and partners closely with engineering leadership to drive strategic initiatives and ensure seamless coordination between research, engineering, and product teams.
Responsibilities
- Drive cross-functional programs to improve developer environments, CI/CD infrastructure, and release processes.
- Coordinate large-scale migrations and platform modernization efforts.
- Partner with teams to measure and improve developer productivity metrics.
- Lead initiatives to integrate AI tools into development workflows.
- Drive programs to establish and achieve reliability targets across training infrastructure and production services.
- Coordinate incident response improvements, post-mortem processes, and on-call rotations.
- Establish metrics and dashboards to track infrastructure health, capacity utilization, and operational excellence.
- Serve as the critical bridge between infrastructure teams, research, and product.
- Consult with stakeholders to understand infrastructure, data, and compute needs.
- Drive alignment on priorities and timelines across teams.
Requirements
- 5+ years of technical program management experience, with a track record of successfully delivering complex infrastructure programs in ML/AI systems or large-scale distributed systems.
- Deep technical understanding of infrastructure systems.
- Excel at creating structure and processes in ambiguous environments.
- Strong stakeholder management skills.
- Comfortable navigating competing priorities and using data to drive technical decisions.
- Experience with developer productivity initiatives, CI/CD systems, or infrastructure scaling.
- Thrive in fast-paced environments and can balance strategic planning with tactical execution.
- Obsessed with reliability, scalability, security, and continuous improvement.
- Passion for supporting internal partners like research.
- Passionate about AI infrastructure and understand the unique challenges of building and operating systems at frontier scale.
- Experience with Kubernetes, cloud platforms (AWS, GCP, Azure), and ML infrastructure (GPU/TPU/Trainium clusters).
- Background working with research teams and translating their needs into concrete technical requirements.
- Experience driving adoption of AI tools to improve engineering productivity.
- Familiarity with observability tooling and practices.
- Bachelor’s degree or an equivalent combination of education, training, and/or experience.