Member of Technical Staff (TPM, Inference)
San Francisco • FullTime
Posted 11d ago
Job Location
San Francisco
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Perplexity is seeking a technical program manager to serve as the crucial link between our model providers, engineering, and product teams, driving the advancement of our core inference platform. This role operates at the nexus of product, engineering, and finance, orchestrating the smooth integration of new models and capacity into production while executing the roadmap for the inference platform itself. The ideal candidate possesses strong technical judgment, excels at coordinating diverse teams and external partners with competing timelines, and is motivated by establishing operating models for functions with limited precedent.
Responsibilities
- Execute the roadmap for the inference platform, including request handling, rate limits, quotas, usage controls, and reliability/observability.
- Act as the connective tissue between model providers and Perplexity's engineering and product teams for model and capacity onboarding, launch readiness, and rollout.
- Drive improvements in latency, throughput, uptime, and cost-efficiency, actively surfacing tradeoffs.
- Manage the operating model for model-release and optimization programs, including day-zero launches, across performance engineering, infrastructure, and product teams.
- Lead cross-functional delivery for inference-stack changes from planning through post-launch validation.
- Develop mechanisms for predictable releases, such as rituals, dashboards, and launch checklists, to ensure low-risk inference releases at scale.
- Partner with GPU capacity and compute teams to align execution decisions with cost, capacity, and vendor constraints.
Requirements
- Strong experience in technical program management or product management within infrastructure, distributed systems, or ML/model-serving products.
- Direct experience with production LLM or ML inference, understanding the nuances of fast, reliable, and cost-effective serving.
- Comfort orchestrating across external partners and internal engineering teams with competing priorities and timelines.
- Experience with data and metrics, with the judgment to surface difficult tradeoffs between latency, throughput, uptime, and cost.
- Ability to thrive in a small, agile team with initiative and a desire for ownership in a less-precedented environment.
- 6+ years of combined technical program management or product management experience.