Product Manager, Inference Platform

Remote San Francisco FullTime

Posted 5mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

Baseten is seeking a Product Manager to define and build the product function for its inference platform, which powers mission-critical AI inference for leading AI companies. This role offers a unique opportunity to shape the future of AI infrastructure, working directly with founders and top engineers. You will own the product surface responsible for making AI model inference fast, reliable, and economical at scale, including autoscaling, traffic routing, failover, and workload scaling across clusters and regions. The ideal candidate is deeply technical, customer-obsessed, and enjoys tackling foundational infrastructure challenges to make AI systems scalable and reliable.

Responsibilities

  • Own workload scaling and placement, from autoscaling to demand to placement policies considering region, compliance, and capacity.
  • Prioritize compliance-bound workloads on sensitive capacity.
  • Ensure production inference reliability by default, guaranteeing requests reach healthy replicas and rolling deploys do not drop traffic.
  • Define region-aware routing with multi-region and active-active failover, including health-aware recovery for replicas.
  • Build the release engine for safe rollouts, including traffic shifting for canary, shadow, and A/B tests, along with warm-ups, drains, and probes.
  • Advance the cost and performance frontier for AI serving across latency, throughput, uptime, and cost efficiency.
  • Drive a measurable reduction in Mean Time To Recovery (MTTR) through self-serve incident management.
  • Set the roadmap for infrastructure and product teams, owning capabilities end-to-end from backend to customer-facing configuration and observability.

Requirements

  • 8+ years of product management experience, with deep expertise in infrastructure, distributed systems, or ML serving.
  • Proficiency in reasoning about scaling, routing, failover, and cost/performance trade-offs at a level respected by staff engineers.
  • Proven experience owning capabilities end-to-end, from backend to user experience.
  • Track record of driving cross-team roadmaps and managing underlying dependencies.
  • Comfort defining new product categories.

Benefits

  • Competitive compensation and meaningful equity
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy with company-wide Winter Break
  • Paid parental leave
  • Fertility and family-building stipend
  • Company-facilitated 401(k)
  • Exposure to a variety of ML startups for learning and networking

About Baseten

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.