Software Engineer- Inference Performance

Remote • San Francisco • FullTime

Posted 2h ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Baseten is seeking inference performance engineers to accelerate and optimize demanding AI workloads. This role involves working across the entire stack, from inference engines and runtimes to scheduling and serving, applying advanced techniques like speculative decoding and KV-cache management. You will analyze performance bottlenecks, improve efficiency, and directly impact the speed and cost-effectiveness of customer models. This position is ideal for individuals who excel in fast-paced startup environments and are passionate about contributing to the LLM inference field.

Responsibilities

  • Implement and productionize inference techniques, including quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, guided generation, and custom scheduling/routing.
  • Profile and optimize inference end-to-end, from kernel launches and memory layout to request scheduling, prefill/decode disaggregation, and cache-aware routing.
  • Improve tokens per GPU-hour, increase utilization, and provide clear latency/throughput/cost tradeoffs.
  • Quickly bring up and tune new model architectures on new hardware.
  • Build benchmarking frameworks to measure real-world performance across various configurations.
  • Contribute to open-source inference engines (vLLM, SGLang, TensorRT-LLM) and collaborate with model, infrastructure, and customer-facing teams.

Requirements

  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or related field.
  • Experience with Python or C++.
  • Familiarity with LLM optimization techniques like quantization and speculative decoding.
  • Strong familiarity with ML libraries such as PyTorch, TensorRT, or TensorRT-LLM.
  • Demonstrated interest and experience in LLMs.
  • Deep understanding of GPU architecture.

Benefits

  • Competitive compensation and meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents (U.S. only).
  • Flexible PTO policy including company-wide Winter Break.
  • Paid parental leave.
  • Fertility and family-building stipend.
  • Company-facilitated 401(k) (U.S. only).
  • Exposure to various ML startups for learning and networking.

About Baseten

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.