ML engineer - API Platform
Remote • San Francisco • FullTime
Posted 5h ago
Job Location
San Francisco
Tech Stack
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Physical Intelligence is building the platform to make foundation models and learning algorithms accessible for powering robots and future physically-actuated devices. As an API Product Engineer, you will develop the product surface that enables other companies to leverage Pi's models similarly to how developers use LLM APIs today. This involves allowing them to bring their own data, fine-tune models, evaluate them, and run low-latency inference within their own environments. You will own this entire experience, transforming capabilities that currently require close team collaboration into a scalable platform designed to support a vast number of robots.
Responsibilities
- Build Pi's model API end-to-end, covering data ingestion, fine-tuning, evaluation, low-latency remote inference, partner tools, and deployment integrations.
- Design scalable systems to support thousands of organizations and potentially millions of robots.
- Architect reliable, multi-tenant infrastructure for rate limiting, isolation, backpressure, versioning, observability, and SLOs.
- Develop a first-class product experience for partner data, from upload through validation, processing, fine-tuning, and evaluation.
- Build and operate low-latency inference systems for Pi models to control real-world robots.
- Collaborate with researchers to translate new model capabilities into stable, usable product surfaces.
- Identify and address recurring friction points based on partner platform usage through improved APIs, tooling, documentation, and abstractions.
- Write high-quality production code that integrates deeply with Pi's existing infrastructure.
- Define the future of the developer platform for general-purpose robotics.
Requirements
- Strong software engineering fundamentals and experience building production systems.
- Deep backend and systems experience across APIs, services, databases, caching, distributed systems, and infrastructure.
- Experience building and scaling developer platforms, particularly for model fine-tuning, inference, or other compute-intensive workloads.
- Understanding of scaling challenges: reliability, latency, multi-tenancy, versioning, observability, and operational complexity.
- Familiarity with machine learning systems for production deployment, serving, and debugging.
- Strong Python skills and comfort working across infrastructure and product boundaries.
- High degree of ownership and ability to take ambiguous problems from inception to a self-sustaining state.