Engineer, Inference
San Francisco, CA • FullTime
Posted 2h ago
About the job
Sierra is seeking a Software Engineer for its Inference team to build and optimize the systems that power customer-facing AI agents. This role focuses on ensuring the speed, reliability, and efficiency of foundation models at scale. You will define the inference architecture, manage serving and routing, handle capacity, and optimize for latency, reliability, and cost. This is a systems-first position at the intersection of distributed infrastructure and AI, ideal for engineers who enjoy complex systems challenges and are eager to apply their expertise to the rapidly evolving field of AI infrastructure.
Responsibilities
- Partner with frontier labs and inference providers for capacity and infrastructure.
- Define and shape Sierra's inference architecture across models, infrastructure, and providers.
- Design and build systems for low latency and high reliability in inference serving.
- Develop and operate self-hosted inference on GPU infrastructure.
- Optimize inference performance in collaboration with the Applied Research team.
- Build and manage across a hybrid inference stack, including Sierra-managed and third-party platforms.
- Collaborate with inference providers to tune engines and infrastructure.
- Contribute to infrastructure supporting the broader model lifecycle and post-training processes.
Requirements
- Deep systems thinking and strong distributed systems fundamentals.
- Experience designing, building, and operating large-scale production systems.
- Strong judgment regarding tradeoffs in latency, reliability, capacity, and cost.
- Experience taking ownership of complex infrastructure from architecture to production operation.
- Excitement for applying systems expertise to AI infrastructure and rapid learning.
- Experience with ML infrastructure, MLOps, or production inference systems.
- Experience serving LLMs or other large models at scale.
- Experience operating self-hosted inference and GPU infrastructure.
- Familiarity with inference frameworks like vLLM or SGLang.
- Experience with post-training infrastructure or inference-performance optimization.
Benefits
- Flexible (unlimited) paid time off
- Medical, dental, and vision benefits for you and your family
- Life insurance and disability benefits