Software Engineer - Model Performance
Remote • San Francisco • FullTime
Posted 2y ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Baseten powers mission-critical inference for leading AI companies, enabling them to bring cutting-edge models into production. We are seeking a Software Engineer focused on ML performance to join our dynamic team. This role is ideal for someone who thrives in a fast-paced startup environment and is eager to make significant contributions to the exciting field of LLM Inference. If you are a backend engineer who thrives on making things faster and is excited about open-source ML models, we look forward to your application.
Responsibilities
- Implement, refine, and productionize techniques like quantization, speculative decoding, KV cache reuse, chunked prefill, and LoRA for ML model inference and infrastructure.
- Debug ML performance issues by deep diving into codebases of libraries such as TensorRT, PyTorch, TensorRT-LLM, vLLM, sglang, and CUDA.
- Apply and scale optimization techniques across a wide range of ML models, particularly large language models.
- Collaborate with a diverse team to design and implement innovative solutions.
- Own projects from conception to production.
Requirements
- Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field.
- Experience with general-purpose programming languages like Python or C++.
- Familiarity with LLM optimization techniques (e.g., quantization, speculative decoding, continuous batching).
- Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM.
- Demonstrated interest and experience in LLMs.
- Deep understanding of GPU architecture.
Benefits
- Competitive compensation, including meaningful equity
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break
- Paid parental leave
- Fertility and family-building stipend
- Company-facilitated 401(k)
- Exposure to a variety of ML startups, offering learning and networking opportunities.