Software Engineer - Model Products
Remote • San Francisco • FullTime
Posted 11mo ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Baseten is seeking a Software Engineer to join their Model Performance team, focusing on the infrastructure that powers hosted API endpoints for cutting-edge open-source models. This role involves working on distributed systems, model serving, and developer experience to ensure models running on the Baseten platform are fast, reliable, and cost-efficient. You will contribute to defining how developers interact with AI models at scale, joining a high-impact team at the intersection of product, model performance, and infrastructure.
Responsibilities
- Design, build, and operate Model APIs with a focus on advanced inference capabilities like structured outputs, tool/function calling, and multi-modal serving.
- Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, and implement custom CUDA operators.
- Tune memory allocation patterns for maximum throughput and optimize communication across multi-GPU setups.
- Productionize performance improvements across runtimes with a deep understanding of their internals, including speculative decoding, guided generation, and custom scheduling algorithms.
- Build comprehensive benchmarking frameworks to measure real-world performance across various configurations.
- Productionize performance improvements across runtimes like TensorRT and TensorRT-LLM, focusing on speculative decoding, quantization, batching, and KV-cache reuse.
- Instrument deep observability (metrics, traces, logs) and build repeatable benchmarks for speed, reliability, and quality.
- Implement platform fundamentals such as API versioning, validation, usage metering, quotas, and authentication.
- Collaborate with other teams to deliver robust and developer-friendly model serving experiences.
Requirements
- 3+ years of experience building and operating distributed systems or large-scale APIs.
- Proven track record of owning low-latency, reliable backend services.
- Strong infrastructure instincts with performance sensibilities, including profiling, tracing, capacity planning, and SLO management.
- Comfortable debugging complex systems, from runtime internals to GPU execution traces.
- Strong written communication skills for producing clear design documents and collaborating across functions.
Benefits
- Competitive compensation and meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employee and dependents.
- Flexible PTO policy with company-wide Winter Break.
- Paid parental leave.
- Fertility and family-building stipend.
- Company-facilitated 401(k).
- Exposure to a variety of ML startups for learning and networking opportunities.