Software Engineer - Model Performance

Remote San Francisco FullTime

Posted 2y ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Baseten powers mission-critical inference for leading AI companies, enabling them to bring cutting-edge models into production. We are seeking a Software Engineer focused on ML performance to join our dynamic team. This role is ideal for someone who thrives in a fast-paced startup environment and is eager to make significant contributions to the exciting field of LLM Inference. If you are a backend engineer who thrives on making things faster and is excited about open-source ML models, we look forward to your application.

Responsibilities

  • Implement, refine, and productionize techniques like quantization, speculative decoding, KV cache reuse, chunked prefill, and LoRA for ML model inference and infrastructure.
  • Debug ML performance issues by deep diving into codebases of libraries such as TensorRT, PyTorch, TensorRT-LLM, vLLM, sglang, and CUDA.
  • Apply and scale optimization techniques across a wide range of ML models, particularly large language models.
  • Collaborate with a diverse team to design and implement innovative solutions.
  • Own projects from conception to production.

Requirements

  • Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field.
  • Experience with general-purpose programming languages like Python or C++.
  • Familiarity with LLM optimization techniques (e.g., quantization, speculative decoding, continuous batching).
  • Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM.
  • Demonstrated interest and experience in LLMs.
  • Deep understanding of GPU architecture.

Benefits

  • Competitive compensation, including meaningful equity
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break
  • Paid parental leave
  • Fertility and family-building stipend
  • Company-facilitated 401(k)
  • Exposure to a variety of ML startups, offering learning and networking opportunities.

About Baseten

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.