Engineering Manager - Inference Performance
$50k - $90k • Remote • San Francisco • FullTime
Posted 6h ago
About the job
Baseten is seeking an Engineering Manager to lead a portion of its Inference Performance team. This team is responsible for accelerating and optimizing demanding AI workloads on GPUs. The role involves managing and developing a team of inference performance engineers, focusing on areas like the inference engine and runtime, including kernels, scheduling, batching, KV-cache management, speculative decoding, and prefill/decode disaggregation. It's a hands-on technical leadership position where you will define the team's direction, solve complex challenges, and build trust through deep engagement with GPU performance, while also focusing on hiring, developing, and supporting your team members. The team's contributions directly impact the speed and efficiency of customer models.
Responsibilities
- Lead, mentor, and grow a team of inference performance engineers through regular 1:1s, feedback, career development, and performance reviews.
- Hire top GPU and inference engineering talent and foster a strong, collaborative team culture.
- Own the technical roadmap and execution for runtime performance work, balancing customer needs, new model launches, and long-term platform investments.
- Review designs, guide profiling and optimization efforts, and help the team understand performance bottlenecks.
- Drive the productionization of inference techniques like quantization, speculative decoding, KV-cache reuse, chunked prefill, and custom scheduling.
- Translate performance improvements into measurable outcomes such as tokens per GPU-hour, utilization, latency, and cost.
- Assist the team in quickly bringing up and tuning new model architectures on new hardware.
- Partner with cross-functional teams (Infrastructure, Inference Platform, Kernels, Model APIs, customer-facing) to set priorities and coordinate launches.
- Establish high standards for engineering quality, benchmarking, operational excellence, and incident response.
Requirements
- Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field.
- Proven experience managing engineers, including hiring, mentoring, providing feedback, and conducting performance reviews.
- Experience leading or closely supporting GPU optimization teams in training, inference, or recommendation systems.
- Strong technical depth in GPU workloads, including a solid understanding of GPU architecture and performance tradeoffs.
- Familiarity with ML libraries such as PyTorch, TensorRT, or TensorRT-LLM.
- A track record of driving roadmaps and successfully shipping complex technical projects with a team.
- Clear written and verbal communication skills, with the ability to align stakeholders across different teams.
Benefits
- Competitive compensation, including meaningful equity
- 100% coverage of medical, dental, and vision insurance for employee and dependents (U.S. only)
- Flexible PTO policy
- Company-wide Winter Break (offices closed from Christmas Eve to New Year's Day)
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k) (U.S. only)
- Exposure to a variety of ML startups, offering learning and networking opportunities