Engineering Manager, Model Inference
Remote • SF Office • FullTime
Posted 2mo ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Abridge is seeking an Engineering Manager to lead and grow our Model Inference team. This role is crucial for ensuring our generative AI-powered products are fast, reliable, and world-class. The Inference team is responsible for the end-to-end technical direction of model serving, from architecting low-latency, high-throughput infrastructure to advancing LLM serving techniques. You will lead a team of AI inference engineers, collaborate with ML Research and the AI Platform, and ensure the systems supporting clinician interactions operate at peak efficiency and reliability.
Responsibilities
- Lead and grow a team of AI inference engineers focused on building and scaling infrastructure for products and APIs.
- Own the technical direction of inference systems, making key decisions on batching, throughput, latency, and GPU utilization.
- Architect and scale inference infrastructure for reliability, efficiency, and observability, and lead incident response.
- Benchmark and eliminate bottlenecks throughout the inference stack.
- Partner with ML Research teams on model optimization, quantization, and deployment.
- Develop APIs for AI inference used by internal teams and external customers.
- Recruit, mentor, and develop engineering talent, establishing team processes, engineering standards, and operational excellence.
- Collaborate with GenAI Platform, Data, and Product teams to plan and execute projects impacting clinicians and patients.
Requirements
- 5+ years of engineering experience, with at least 1 year in a technical leadership or management role.
- Deep, hands-on experience with ML systems and inference frameworks (e.g., PyTorch, TensorRT, vLLM, TensorFlow).
- Strong understanding of LLM architecture (e.g., Multi-Head Attention, Multi/Grouped-Query Attention, common transformer components).
- Experience with inference optimizations (e.g., batching, quantization, kernel fusion, FlashAttention).
- Familiarity with GPU characteristics, roofline models, and performance analysis.
- Experience deploying reliable, distributed, real-time systems at scale.
- Experience with parallelism strategies: tensor parallelism, pipeline parallelism, expert parallelism.
- Skilled at hiring and mentorship, with a demonstrated track record of helping engineers grow.
- Strong technical communication and cross-functional collaboration skills.
- Comfortable giving constructive feedback on technical designs and code reviews.
- Experience thriving in a fast-growing startup environment with urgency and focus.
Benefits
- 14 paid holidays
- Flexible PTO for salaried employees
- Accrued time off for hourly employees
- Comprehensive Medical, Dental, and Vision coverage for full-time employees and their families
- Generous HSA Contribution if choosing a High Deductible Health Plan
- Generous paid parental leave for all full-time employees
- Family Forming Benefits