Software Engineer, Inference
$30k - $60k • Remote • Redwood City, CA • FullTime
Posted 1mo ago
Job Location
Redwood City, CA
Tech Stack
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Luma is seeking a Software Engineer to own the serving of their models. This role involves integrating new architectures into the inference engine, scaling deployments across thousands of machines, and optimizing GPU fleet utilization while meeting internal service level objectives (SLOs). The work focuses on large-scale inference systems, including scheduling, fleet management, deployment pipelines, and reliability across various clusters and hardware providers. This position is ideal for a strong systems engineer experienced with model serving and Kubernetes at scale, rather than pure modeling.
Responsibilities
- Integrate new model architectures into the inference engine.
- Collaborate with research, engineering, and infrastructure teams to optimize model efficiency and deployments.
- Develop internal tooling to measure, profile, and track inference jobs and workflows.
- Automate, test, and maintain inference services for maximum uptime and reliability.
- Manage and optimize inference workloads across clusters and hardware providers, scaling deployments across thousands of machines.
- Build scheduling systems for optimal GPU resource utilization and SLO adherence.
- Maintain CI/CD for model checkpoints and SDKs.
Requirements
- Strong Python and system-architecture skills.
- Experience deploying models with frameworks like PyTorch, Hugging Face, vLLM, SGLang, or TensorRT-LLM.
- Experience with queues, scheduling, traffic control, and fleet management at scale.
- Proficiency with Linux, Docker, and Kubernetes, including orchestration, deployment, and scheduling.
- Familiarity with Redis and S3-compatible storage.