Senior ML Systems Engineer, Inference
$150k - $220k • Remote • Remote - USA • FullTime
Posted 6h ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Machine Learning Engineer
About the job
Runpod is seeking a Senior ML Systems Engineer, Inference to lead efforts in making their AI Developer Cloud the premier platform for LLM inference. This role focuses on optimizing LLM serving performance end-to-end, encompassing measurement, understanding, and improvement across various models, hardware, and workloads. The successful candidate will be responsible for hands-on engineering to identify and resolve performance bottlenecks, translating these improvements into reliable production systems that directly impact customer latency and cost.
Responsibilities
- Define and build tooling for rigorous and repeatable inference performance measurements (throughput, time to first token, inter-token latency, cost per token).
- Profile and diagnose performance issues across the serving stack, from scheduling to kernels and interconnect.
- Enhance serving efficiency for large, state-of-the-art models on single and multi-node GPU deployments.
- Develop production-ready runtimes, configurations, and defaults based on performance findings.
- Collaborate with product and infrastructure teams to define inference offerings.
- Stay abreast of the inference ecosystem, including open-source developments, and determine adoption, building, or contribution strategies.
- Identify and fix bottlenecks in the serving engine/runtime when configuration tuning is insufficient.
Requirements
- 5+ years of professional system engineering experience.
- Deep, hands-on experience with vLLM, SGLang, or comparable serving engines in production or at scale.
- Strong software engineering skills in Python, with experience in large, performance-critical codebases.
- Solid understanding of LLM inference performance drivers (batching, memory, parallelism, latency/throughput trade-offs).
- Experience with modern inference optimization techniques (quantization, speculative decoding, distributed serving).
- Rigor in benchmarking and performance analysis, including proficiency with GPU profiling tools.
- Ability to clearly explain results in writing and translate them into actionable decisions.
Benefits
- Competitive base pay ranging from $150,000 - $220,000.
- Meaningful equity and stock options.
- Generous medical, dental, and vision plans.
- Flexible PTO.
- Remote work opportunities.
- $1,200 Home Office & Equipment Stipend.