Staff Applied AI Inference Engineer

$215k - $260k San Francisco, CA - US FullTime

Posted 5d ago

Job Location

San Francisco, CA - US

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Crusoe is seeking a Staff Applied AI Inference Engineer to accelerate the abundance of energy and intelligence by optimizing large language models for production environments. This role involves owning the inference stack end-to-end, from profiling costs and implementing modern optimization techniques to deep dives into serving code when defaults are insufficient. The work is applied, focusing on real-world deployments with specific models, traffic patterns, latency targets, and cost constraints. You will collaborate with customer engineering teams to tailor deployments, transition workloads from proof-of-concept to fully monitored production services, and ensure engineered gains benefit end-users. This is a hands-on engineering position requiring coding, profiling, and low-level optimization, with a customer-facing component involving product and technical solutions work.

Responsibilities

  • Bring current inference techniques into production and refine them.
  • Design and optimize serving architectures, including prefill and decode disaggregation, request routing, and related approaches.
  • Work down into the serving stack, from frameworks like vLLM and SGLang to CUDA kernels, profiling and analyzing performance to fix problems.
  • Adapt and scale optimization methods across various ML models, with an emphasis on large language models.
  • Profile and tune deployments against targets for latency, throughput, and cost, ensuring dependability under real traffic.
  • Tailor deployments to customer models and constraints, partnering with their engineering teams to move workloads from proof-of-concept to production.
  • Build and support software and product features around the inference stack in production using general-purpose languages, preferably Python.
  • Experiment quickly by shaping fuzzy goals into clear specs and focused proofs of concept, running fast experiments, and shipping well-tested results.
  • Own delivery end-to-end, from experiment to production optimization, maintaining focus on performance goals and drafting product requirement documents.
  • Work through ambiguity, make sound decisions on tradeoffs and tooling, and steer away from unnecessary complexity.
  • Take pride and ownership in work, hold yourself accountable, and expect the same from colleagues.

Requirements

  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
  • Hands-on experience shipping code in production with one or more general-purpose languages, such as Python or C++, with a strong preference for Python.
  • Familiarity with methods for optimizing LLMs for high throughput / low latency inference.
  • Comfort with modern LLM serving frameworks such as vLLM or SGLang, and with profiling and analyzing performance down to the kernel level.
  • A firm grasp of how GPUs are built and how they behave.
  • Clear interest and hands-on experience with large language models.
  • Working knowledge of AI/ML pipelines and the full path of developing and deploying ML models.
  • Strong communication skills, particularly when explaining hard technical topics to customers and teammates.

Benefits

  • Competitive compensation and equity packages
  • Restricted Stock Units
  • Paid time off, paid holidays & leave of absence programs
  • Comprehensive health, dental & vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance, short-term and long-term disability
  • Professional development & tuition reimbursement
  • Mental health & wellness support
  • Commuter benefits (parking & transit)
  • Cell phone stipend

About crusoe

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.