Software Engineer- Inference Platform
Remote • San Francisco • FullTime
Posted 2h ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
Baseten is seeking distributed systems engineers and product-minded generalists to build the distributed runtime that powers large-scale LLM inference. Our inference platform enables customers to deploy and operate cutting-edge models with industry-leading performance, scalability, and reliability. You will work across the stack, from the developer experience for model deployment to the systems that orchestrate deployments on Kubernetes and route traffic efficiently, ensuring every model on our platform is fast, reliable, and cost-efficient. This role is ideal for engineers who enjoy owning systems in production, solving hard integration problems, and making complex infrastructure simple and reliable for users.
Responsibilities
- Build infrastructure and orchestration systems for large-scale distributed LLM inference, including routing, autoscaling, scheduling, and runtime management.
- Design, build, and operate Model APIs with a focus on advanced inference capabilities like structured outputs, tool/function calling, and multimodal serving.
- Implement platform fundamentals such as API versioning, validation, usage metering, quotas, and authentication.
- Instrument deep observability (metrics, traces, logs) and build repeatable benchmarks for speed, reliability, and quality.
- Debug and harden complex production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads.
- Partner with Inference Performance engineers to make new optimizations broadly available and easy to configure.
- Own projects end-to-end, from architecture through deployment, monitoring, and iteration on customer feedback, making thoughtful tradeoffs between performance, reliability, operational simplicity, and developer experience.
Requirements
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 3+ years building and operating distributed systems, backend infrastructure, or large-scale APIs where reliability, latency, and scale are critical.
- Proven track record of owning low-latency, reliable backend services, including rate limiting, auth, quotas, metering, and migrations.
- Infrastructure instincts with a feel for performance, including profiling, tracing, capacity planning, and SLO management.
- Comfort debugging performance and reliability issues across multiple layers of the stack.
- Strong sense of developer experience, considering how systems are used.
- Eagerness to learn new languages, frameworks, and systems, with an interest in inference engineering.
- Excellent written communication and collaboration skills, including writing clear design docs.
Benefits
- Competitive compensation, including meaningful equity
- 100% coverage of medical, dental, and vision insurance for employee and dependents (U.S. only)
- Flexible PTO policy including company-wide Winter Break
- Paid parental leave
- Fertility and family-building stipend
- Company-facilitated 401(k) (U.S. only)
- Exposure to a variety of ML startups, offering learning and networking opportunities