Senior ML Systems Engineer, Inference

$150k - $220k • Remote • Remote - USA • FullTime

Posted 6h ago

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Machine Learning Engineer

About the job

Runpod is seeking a Senior ML Systems Engineer, Inference to lead efforts in making their AI Developer Cloud the premier platform for LLM inference. This role focuses on optimizing LLM serving performance end-to-end, encompassing measurement, understanding, and improvement across various models, hardware, and workloads. The successful candidate will be responsible for hands-on engineering to identify and resolve performance bottlenecks, translating these improvements into reliable production systems that directly impact customer latency and cost.

Responsibilities

  • Define and build tooling for rigorous and repeatable inference performance measurements (throughput, time to first token, inter-token latency, cost per token).
  • Profile and diagnose performance issues across the serving stack, from scheduling to kernels and interconnect.
  • Enhance serving efficiency for large, state-of-the-art models on single and multi-node GPU deployments.
  • Develop production-ready runtimes, configurations, and defaults based on performance findings.
  • Collaborate with product and infrastructure teams to define inference offerings.
  • Stay abreast of the inference ecosystem, including open-source developments, and determine adoption, building, or contribution strategies.
  • Identify and fix bottlenecks in the serving engine/runtime when configuration tuning is insufficient.

Requirements

  • 5+ years of professional system engineering experience.
  • Deep, hands-on experience with vLLM, SGLang, or comparable serving engines in production or at scale.
  • Strong software engineering skills in Python, with experience in large, performance-critical codebases.
  • Solid understanding of LLM inference performance drivers (batching, memory, parallelism, latency/throughput trade-offs).
  • Experience with modern inference optimization techniques (quantization, speculative decoding, distributed serving).
  • Rigor in benchmarking and performance analysis, including proficiency with GPU profiling tools.
  • Ability to clearly explain results in writing and translate them into actionable decisions.

Benefits

  • Competitive base pay ranging from $150,000 - $220,000.
  • Meaningful equity and stock options.
  • Generous medical, dental, and vision plans.
  • Flexible PTO.
  • Remote work opportunities.
  • $1,200 Home Office & Equipment Stipend.

About Runpod

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.