Senior Manager, Engineering - AI Inference
$250k - $300k • San Francisco, CA - US • FullTime
Posted 10h ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Crusoe is seeking a Senior Engineering Manager to lead a team focused on optimizing large language models for speed, cost, and reliability in production. This role requires a blend of people management and deep technical involvement in the end-to-end inference stack, including performance profiling, implementing optimization techniques, and working directly with serving code. The work is practical, focusing on real-world customer deployments with diverse models, traffic patterns, and constraints. You will collaborate with customer engineering teams to tailor solutions, transition workloads from proof-of-concept to production, and ensure performance gains deliver tangible business value. This leadership position involves people management, technical direction, hands-on coding, profiling, and customer-facing responsibilities.
Responsibilities
- Implement and refine current inference techniques in production environments.
- Design and optimize serving architectures, including disaggregation strategies and request routing.
- Analyze and optimize the serving stack, from frameworks like vLLM and SGLang down to CUDA kernels.
- Adapt and scale optimization methods for various ML models, with a focus on LLMs.
- Profile and tune deployments for latency, throughput, and cost targets, ensuring reliability under real traffic.
- Customize deployments for specific customer models and constraints, partnering with their engineering teams.
- Develop and support software and product features for the inference stack in production, preferably using Python.
- Rapidly experiment with fuzzy goals, define clear specifications, and ship well-tested results.
- Lead team delivery from early experiments to production optimizations, setting performance goals, and contributing to technical strategy.
- Navigate ambiguity, make informed decisions on tradeoffs and tooling, and avoid unnecessary complexity.
- Take ownership and accountability for work, and foster the same in team members.
Requirements
- 2+ years of experience managing and leading engineering teams in high-performance or ML-focused environments.
- Strong hands-on experience in software engineering, low-level optimization, or ML infrastructure, with a desire to stay close to code.
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
- Shipped production code in general-purpose languages like Python or C++, with a preference for Python.
- Familiarity with methods for optimizing LLMs for high throughput and low latency inference.
- Experience with LLM serving frameworks (e.g., vLLM, SGLang) and performance analysis down to the kernel level.
- Understanding of GPU architecture and behavior.
- Hands-on experience with large language models.
- Working knowledge of AI/ML pipelines and the end-to-end ML model development and deployment process.
- Strong communication skills, especially for explaining complex technical topics.
Benefits
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off