Staff + Sr. Software Engineer, Cloud Inference
San Francisco, CA
Posted 16d ago
About the job
The Cloud Inference team is responsible for scaling and optimizing Claude to serve massive audiences of developers and enterprise companies across AWS, GCP, Azure, and other cloud service providers. This role involves owning the end-to-end product of Claude on each cloud platform, from API integration and intelligent request routing to inference execution, capacity management, and daily operations. Engineers on this team have a high leverage impact, driving multiple revenue streams while optimizing compute resources. The complexity of managing inference efficiently across diverse providers with different hardware, networking, and operational models is growing significantly. We are seeking product-minded backend engineers who can navigate these platform differences, design cross-provider services and abstractions, and make architectural decisions to ensure reliability and cost-effectiveness at massive scale. Your work will directly contribute to increasing service scale, accelerating the launch of new models and features, and ensuring LLMs meet rigorous safety, performance, and security standards.
Responsibilities
- Design, build, and own backend services and infrastructure for serving Claude across multiple CSPs, considering differences in compute hardware, networking, APIs, and operational models.
- Collaborate cross-functionally with internal inference, product API, systems, and security teams, and with CSP partners to deploy the full serving stack on new cloud platforms, resolve operational issues, and influence provider roadmaps.
- Build and evolve CI/CD automation systems, including validation and deployment pipelines, for reliably shipping new model versions to millions of users across cloud platforms without regressions.
- Design interfaces and tooling abstractions across CSPs to enable cost-effective inference management, scale across providers, and reduce per-platform complexity.
- Contribute to capacity planning, autoscaling, and workload routing strategies to match supply with demand and direct requests to the most cost-effective accelerator and region.
- Analyze observability data across providers to identify performance bottlenecks, cost anomalies, and regressions, and drive remediation based on real-world production workloads.
Requirements
- Significant software engineering experience with a strong background in high-performance, large-scale distributed systems serving millions of users.
- Experience building or operating services on at least one major cloud platform (AWS, GCP, or Azure), with exposure to Kubernetes, Infrastructure as Code, or container orchestration.
- Curiosity about LLM serving; prior inference or ML experience is not required.
- Ability to thrive in cross-functional collaboration with both internal teams and external partners.
- Experience working with external partners to align goals and deliver impact.
- Fast learner capable of quickly ramping up on new technologies, hardware platforms, and provider ecosystems.
- Highly autonomous with the ability to take ownership of problems end-to-end, including work that falls outside the defined job description.
- Proficiency in Python or Rust.