Software Engineer, Inference

$30k - $60k Remote Redwood City, CA FullTime

Posted 1mo ago

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Luma is seeking a Software Engineer to own the serving of their models. This role involves integrating new architectures into the inference engine, scaling deployments across thousands of machines, and optimizing GPU fleet utilization while meeting internal service level objectives (SLOs). The work focuses on large-scale inference systems, including scheduling, fleet management, deployment pipelines, and reliability across various clusters and hardware providers. This position is ideal for a strong systems engineer experienced with model serving and Kubernetes at scale, rather than pure modeling.

Responsibilities

  • Integrate new model architectures into the inference engine.
  • Collaborate with research, engineering, and infrastructure teams to optimize model efficiency and deployments.
  • Develop internal tooling to measure, profile, and track inference jobs and workflows.
  • Automate, test, and maintain inference services for maximum uptime and reliability.
  • Manage and optimize inference workloads across clusters and hardware providers, scaling deployments across thousands of machines.
  • Build scheduling systems for optimal GPU resource utilization and SLO adherence.
  • Maintain CI/CD for model checkpoints and SDKs.

Requirements

  • Strong Python and system-architecture skills.
  • Experience deploying models with frameworks like PyTorch, Hugging Face, vLLM, SGLang, or TensorRT-LLM.
  • Experience with queues, scheduling, traffic control, and fleet management at scale.
  • Proficiency with Linux, Docker, and Kubernetes, including orchestration, deployment, and scheduling.
  • Familiarity with Redis and S3-compatible storage.

About lumalabs

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.