Software Engineer - Model Products

Remote San Francisco FullTime

Posted 11mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Baseten is seeking a Software Engineer to join their Model Performance team, focusing on the infrastructure that powers hosted API endpoints for cutting-edge open-source models. This role involves working on distributed systems, model serving, and developer experience to ensure models running on the Baseten platform are fast, reliable, and cost-efficient. You will contribute to defining how developers interact with AI models at scale, joining a high-impact team at the intersection of product, model performance, and infrastructure.

Responsibilities

  • Design, build, and operate Model APIs with a focus on advanced inference capabilities like structured outputs, tool/function calling, and multi-modal serving.
  • Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, and implement custom CUDA operators.
  • Tune memory allocation patterns for maximum throughput and optimize communication across multi-GPU setups.
  • Productionize performance improvements across runtimes with a deep understanding of their internals, including speculative decoding, guided generation, and custom scheduling algorithms.
  • Build comprehensive benchmarking frameworks to measure real-world performance across various configurations.
  • Productionize performance improvements across runtimes like TensorRT and TensorRT-LLM, focusing on speculative decoding, quantization, batching, and KV-cache reuse.
  • Instrument deep observability (metrics, traces, logs) and build repeatable benchmarks for speed, reliability, and quality.
  • Implement platform fundamentals such as API versioning, validation, usage metering, quotas, and authentication.
  • Collaborate with other teams to deliver robust and developer-friendly model serving experiences.

Requirements

  • 3+ years of experience building and operating distributed systems or large-scale APIs.
  • Proven track record of owning low-latency, reliable backend services.
  • Strong infrastructure instincts with performance sensibilities, including profiling, tracing, capacity planning, and SLO management.
  • Comfortable debugging complex systems, from runtime internals to GPU execution traces.
  • Strong written communication skills for producing clear design documents and collaborating across functions.

Benefits

  • Competitive compensation and meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents.
  • Flexible PTO policy with company-wide Winter Break.
  • Paid parental leave.
  • Fertility and family-building stipend.
  • Company-facilitated 401(k).
  • Exposure to a variety of ML startups for learning and networking opportunities.

About Baseten

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.