Machine Learning Infrastructure Engineer, Model Inference
Remote • SF Office • FullTime
Posted 11mo ago
Job Location
SF Office
Tech Stack
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
Abridge is seeking an ML Infrastructure Engineer, Model Inference to build and optimize the core inference infrastructure powering their machine learning models. This role is crucial for enhancing the scalability, efficiency, and performance of Abridge's AI-driven healthcare solutions. The engineer will collaborate with Infrastructure and Research teams to build, deploy, optimize, and orchestrate AI models, working on a platform that transforms patient-clinician conversations into structured clinical notes in real-time.
Responsibilities
- Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training.
- Develop, optimize, and maintain ML model serving infrastructure for high performance and low latency.
- Collaborate with ML and product teams to scale backend infrastructure for AI-driven products, focusing on model deployment, throughput, and compute efficiency.
- Optimize compute-heavy workflows and enhance GPU utilization for ML workloads.
- Build a robust model API orchestration system.
- Collaborate with leadership to define and implement strategies for scaling infrastructure.
Requirements
- 2+ years of experience building and deploying machine learning models in production.
- Deep understanding of container orchestration and distributed systems architecture.
- Expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management.
- Experience developing APIs and managing distributed systems for batch and real-time workloads.
- Excellent communication skills, with the ability to interface between research and product engineering.
Benefits
- 14 paid holidays
- Flexible PTO for salaried employees
- Accrued time off for hourly employees
- Comprehensive Medical, Dental, and Vision coverage