Engineering Manager, Inference Infrastructure
San Francisco, CA | New York City, NY | Seattle, WA
Posted 13d ago
Job Location
San Francisco, CA | New York City, NY | Seattle, WA
Tech Stack
Remote Work Policy
On-site
Categories
AI Infrastructure Engineer
About the job
Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. This role leads a team of ML platform, infrastructure, and distributed-systems engineers responsible for the critical control plane that manages Anthropic's inference fleet. This involves making key decisions about request routing, capacity allocation, and system performance to meet throughput, reliability, and latency constraints. The team designs placement and load-balancing algorithms, builds quantitative models for demand and capacity, and optimizes latency across various system boundaries. The Engineering Manager will own the technical roadmap, drive quantitative modeling habits, set technical strategy for evolving the control plane, and ensure the operational health of the inference request path.
Responsibilities
- Own the technical roadmap for inference fleet coordination, including traffic routing, capacity management, cache placement, demand response, and control plane/inference engine synchronization.
- Partner with product, inference engine, performance, and capacity teams to identify and implement throughput, latency, utilization, and cost improvements.
- Foster a habit of quantitative modeling to measure the impact of changes and predict outcomes before deployment.
- Set technical strategy for the evolution of the control plane across heterogeneous hardware, multiple cloud providers, and various serving surfaces.
- Manage the group's operational backbone, including on-call rotations, incident response, postmortem reviews, and deploy safety.
- Create clarity at the intersection of API surfaces, inference engines, capacity planning, and cloud deployment teams.
- Develop and retain existing team members, and hire new talent with a high technical bar.
- Coach engineers through a roadmap with shifting priorities.
- Shape team structure as scope grows, defining problem area boundaries and developing leads.
- Address critical path blockers and synthesize design debates when necessary.
Requirements
- Engineering management experience leading teams on critical-path production infrastructure at scale.
- Deep systems background in areas like load balancing, scheduling, cluster orchestration, autoscaling, distributed state, or high-performance networking.
- Ability to make architectural calls for large-scale fleet coordination and evaluate candidates with kernel/framework-level expertise.
- Experience shipping performance or efficiency improvements in large-scale systems with measurable impact (including cost).
- Experience running production infrastructure with operational responsibilities (on-call, incident response, capacity planning, deploy discipline).
- Results-oriented and impact-driven approach, comfortable balancing competing priorities like throughput, latency, cost, stability, and velocity.
- Ability to build strong cross-team relationships in a seam role.
- Curiosity about machine learning systems and a desire to learn how transformer inference impacts systems problems.