Senior Staff Applied AI Inference Engineer
$250k - $300k • San Francisco, CA - US • FullTime
Posted 15h ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Crusoe is seeking a Senior Staff Applied AI Inference Engineer to accelerate the abundance of energy and intelligence by making large language models run faster, cheaper, and more reliably in production. This role involves owning the inference stack end-to-end, from profiling performance bottlenecks to implementing modern optimization techniques and diving deep into serving code when necessary. The work is applied, focusing on real-world deployments with specific models, traffic patterns, latency targets, and cost constraints. You will collaborate with customer engineering teams to tailor deployments, transition workloads from proof-of-concept to fully monitored production services, and ensure engineered gains are realized by users. This is a hands-on engineering position requiring coding, profiling, and low-level optimization, with a customer-facing component involving product and technical solutions work.
Responsibilities
- Bring current inference techniques into production and refine them.
- Design and optimize serving architectures, including prefill and decode disaggregation, request routing, and related approaches.
- Work down into the serving stack, from frameworks like vLLM and SGLang to CUDA kernels, profiling and analyzing performance to fix problems.
- Adapt and scale optimization methods across many kinds of ML models, with an emphasis on large language models.
- Profile and tune deployments against clear targets for latency, throughput, and cost, ensuring dependability under real traffic.
- Tailor deployments to each customer's models and constraints, partnering with their engineering teams to move workloads from proof of concept to live production.
- Build and support software and product features around the inference stack in a production setting, using general-purpose languages like Python.
- Experiment quickly by shaping fuzzy goals into clear specs and focused proofs of concept, running fast experiments, and shipping well-tested results.
- Own delivery end to end, from experiment to production optimization, keeping performance goals and specs in mind, and drafting features and product requirement documents.
- Work through ambiguity and make sound calls on tradeoffs and tooling, steering away from unnecessary complexity.
- Take pride and ownership in work, hold yourself accountable, and look for the same from colleagues.
Requirements
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
- Hands-on experience shipping code in production with one or more general-purpose languages, such as Python or C++, with a strong preference for Python.
- Familiarity with methods for optimizing LLMs for high throughput / low latency inference.
- Comfort with modern LLM serving frameworks such as vLLM or SGLang, and with profiling and analyzing performance down to the kernel level.
- A firm grasp of how GPUs are built and how they behave.
- Clear interest and hands-on experience with large language models.
- A working knowledge of AI/ML pipelines and the full path of developing and deploying ML models.
- Strong communication skills, particularly when explaining hard technical topics to customers and teammates.
Benefits
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off, paid holidays & leave of absence programs
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend