Engineer, Inference

San Francisco, CA • FullTime

Posted 2h ago

Job Location

San Francisco, CA

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

Sierra is seeking a Software Engineer for its Inference team to build and optimize the systems that power customer-facing AI agents. This role focuses on ensuring the speed, reliability, and efficiency of foundation models at scale. You will define the inference architecture, manage serving and routing, handle capacity, and optimize for latency, reliability, and cost. This is a systems-first position at the intersection of distributed infrastructure and AI, ideal for engineers who enjoy complex systems challenges and are eager to apply their expertise to the rapidly evolving field of AI infrastructure.

Responsibilities

  • Partner with frontier labs and inference providers for capacity and infrastructure.
  • Define and shape Sierra's inference architecture across models, infrastructure, and providers.
  • Design and build systems for low latency and high reliability in inference serving.
  • Develop and operate self-hosted inference on GPU infrastructure.
  • Optimize inference performance in collaboration with the Applied Research team.
  • Build and manage across a hybrid inference stack, including Sierra-managed and third-party platforms.
  • Collaborate with inference providers to tune engines and infrastructure.
  • Contribute to infrastructure supporting the broader model lifecycle and post-training processes.

Requirements

  • Deep systems thinking and strong distributed systems fundamentals.
  • Experience designing, building, and operating large-scale production systems.
  • Strong judgment regarding tradeoffs in latency, reliability, capacity, and cost.
  • Experience taking ownership of complex infrastructure from architecture to production operation.
  • Excitement for applying systems expertise to AI infrastructure and rapid learning.
  • Experience with ML infrastructure, MLOps, or production inference systems.
  • Experience serving LLMs or other large models at scale.
  • Experience operating self-hosted inference and GPU infrastructure.
  • Familiarity with inference frameworks like vLLM or SGLang.
  • Experience with post-training infrastructure or inference-performance optimization.

Benefits

  • Flexible (unlimited) paid time off
  • Medical, dental, and vision benefits for you and your family
  • Life insurance and disability benefits

About sierra.ai

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.