ML Systems Engineer, Robotics

$249k - $311k Remote San Francisco, CA

Posted 2mo ago

Job Location

San Francisco, CA

Tech Stack

Remote Work Policy

Fully remote

Categories

Machine Learning Engineer

About the job

Scale's Physical AI business unit is focused on solving data bottlenecks in Robotics, Autonomous Vehicles, and Computer Vision. This role involves applied research and developing ML pipelines for processing, training, and fine-tuning data collected by Scale, with an emphasis on optimizing algorithms and pipelines for efficient GPU execution in the cloud. You will advance research, shape Scale's offerings, and expand the frontier of data and model evaluation for Physical AI. As an ML Systems Engineer, you will design and build platforms for scalable, reliable, and efficient serving of foundation models tailored for physical agents, powering both internal research and external customer use cases.

Responsibilities

  • Build and scale fault-tolerant, high-performance systems for serving robotics and foundation models at scale, ensuring low latency for real-time applications.
  • Develop an internal platform to enable model capability discovery and faster iteration cycles for robotics research teams.
  • Collaborate with Robotics researchers and Computer Vision engineers to integrate and optimize models for production and research environments.
  • Conduct architecture and design reviews to ensure system scalability, reliability, and security best practices.
  • Develop monitoring and observability solutions for system health and real-time performance tracking of model inference.
  • Own projects end-to-end, from requirements gathering to implementation, in a fast-paced, cross-functional environment.

Requirements

  • 4+ years of experience building large-scale, high-performance backend systems, with deep experience in machine learning infrastructure.
  • Deep experience optimizing computer vision and other machine learning algorithms for cloud environments, including GPU-level algorithm optimizations (e.g., CUDA, kernel tuning).
  • Strong skills in one or more systems-level languages (e.g., Python, Go, Rust, C++).
  • Deep understanding of serving and routing fundamentals (e.g., rate limiting, load balancing, compute budgets, concurrency) for data-intensive applications.
  • Experience with containers (Docker), orchestration (Kubernetes), and cloud providers (AWS/GCP).
  • Familiarity with infrastructure as code (e.g., Terraform).
  • Proven ability to solve complex problems and work independently in fast-moving environments.

Benefits

  • Base salary
  • Equity
  • Comprehensive health, dental and vision coverage
  • Retirement benefits
  • Learning and development stipend
  • Generous PTO
  • Commuter stipend

About Scale AI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.