Senior Manager, Engineering - AI Inference

$250k - $300k San Francisco, CA - US FullTime

Posted 10h ago

Job Location

San Francisco, CA - US

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Crusoe is seeking a Senior Engineering Manager to lead a team focused on optimizing large language models for speed, cost, and reliability in production. This role requires a blend of people management and deep technical involvement in the end-to-end inference stack, including performance profiling, implementing optimization techniques, and working directly with serving code. The work is practical, focusing on real-world customer deployments with diverse models, traffic patterns, and constraints. You will collaborate with customer engineering teams to tailor solutions, transition workloads from proof-of-concept to production, and ensure performance gains deliver tangible business value. This leadership position involves people management, technical direction, hands-on coding, profiling, and customer-facing responsibilities.

Responsibilities

  • Implement and refine current inference techniques in production environments.
  • Design and optimize serving architectures, including disaggregation strategies and request routing.
  • Analyze and optimize the serving stack, from frameworks like vLLM and SGLang down to CUDA kernels.
  • Adapt and scale optimization methods for various ML models, with a focus on LLMs.
  • Profile and tune deployments for latency, throughput, and cost targets, ensuring reliability under real traffic.
  • Customize deployments for specific customer models and constraints, partnering with their engineering teams.
  • Develop and support software and product features for the inference stack in production, preferably using Python.
  • Rapidly experiment with fuzzy goals, define clear specifications, and ship well-tested results.
  • Lead team delivery from early experiments to production optimizations, setting performance goals, and contributing to technical strategy.
  • Navigate ambiguity, make informed decisions on tradeoffs and tooling, and avoid unnecessary complexity.
  • Take ownership and accountability for work, and foster the same in team members.

Requirements

  • 2+ years of experience managing and leading engineering teams in high-performance or ML-focused environments.
  • Strong hands-on experience in software engineering, low-level optimization, or ML infrastructure, with a desire to stay close to code.
  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
  • Shipped production code in general-purpose languages like Python or C++, with a preference for Python.
  • Familiarity with methods for optimizing LLMs for high throughput and low latency inference.
  • Experience with LLM serving frameworks (e.g., vLLM, SGLang) and performance analysis down to the kernel level.
  • Understanding of GPU architecture and behavior.
  • Hands-on experience with large language models.
  • Working knowledge of AI/ML pipelines and the end-to-end ML model development and deployment process.
  • Strong communication skills, especially for explaining complex technical topics.

Benefits

  • Competitive compensation and equity packages
  • Restricted Stock Units
  • Paid time off

About crusoe

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.