Software Engineer - Baseten Inference Stack

Remote San Francisco FullTime

Posted 3mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Baseten is seeking a Software Engineer to join their Inference Stack team. This team builds the distributed runtime that powers large-scale LLM inference across Baseten's platform, operating at the intersection of distributed systems, model performance, infrastructure, and developer experience. The role involves working across the entire stack, from customer-facing deployment tools and feature libraries to the underlying systems for orchestrating Kubernetes deployments and routing traffic. This is an ideal opportunity for engineers who thrive on owning production systems, solving complex integration challenges, and simplifying intricate infrastructure for users.

Responsibilities

  • Develop infrastructure and orchestration systems for large-scale distributed LLM inference.
  • Work across the stack, from customer-facing features to low-level infrastructure.
  • Build platform capabilities for routing, autoscaling, scheduling, observability, and runtime management.
  • Enhance the reliability, scalability, and usability of the inference stack.
  • Collaborate with Model Performance engineers on inference optimizations.
  • Define best practices for testing, release automation, benchmarking, and operational excellence.
  • Debug complex production systems involving Kubernetes, distributed runtimes, networking, and GPU workloads.
  • Make engineering tradeoffs balancing performance, reliability, simplicity, and developer experience.
  • Own projects end-to-end, including architecture, implementation, deployment, monitoring, and iteration.

Requirements

  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, or related field.
  • Strong background in distributed systems, backend infrastructure, or platform engineering.
  • Experience building and operating production systems where reliability, latency, and scale are critical.
  • Strong focus on developer experience.
  • Willingness to learn new languages, frameworks, and systems.
  • Ability to debug complex, multi-layered systems.
  • Genuine interest in inference engineering.
  • Excellent communication and collaboration skills.

Benefits

  • Competitive compensation and meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents.
  • Flexible PTO policy with company-wide Winter Break.
  • Paid parental leave.
  • Fertility and family-building stipend.
  • Company-facilitated 401(k).
  • Exposure to ML startups for learning and networking.

About Baseten

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.