Staff Applied AI Inference Engineer
$215k - $260k • San Francisco, CA - US • FullTime
Posted 5d ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Crusoe is seeking a Staff Applied AI Inference Engineer to accelerate the abundance of energy and intelligence by optimizing large language models for production environments. This role involves owning the inference stack end-to-end, from profiling costs and implementing modern optimization techniques to deep dives into serving code when defaults are insufficient. The work is applied, focusing on real-world deployments with specific models, traffic patterns, latency targets, and cost constraints. You will collaborate with customer engineering teams to tailor deployments, transition workloads from proof-of-concept to fully monitored production services, and ensure engineered gains benefit end-users. This is a hands-on engineering position requiring coding, profiling, and low-level optimization, with a customer-facing component involving product and technical solutions work.
Responsibilities
- Bring current inference techniques into production and refine them.
- Design and optimize serving architectures, including prefill and decode disaggregation, request routing, and related approaches.
- Work down into the serving stack, from frameworks like vLLM and SGLang to CUDA kernels, profiling and analyzing performance to fix problems.
- Adapt and scale optimization methods across various ML models, with an emphasis on large language models.
- Profile and tune deployments against targets for latency, throughput, and cost, ensuring dependability under real traffic.
- Tailor deployments to customer models and constraints, partnering with their engineering teams to move workloads from proof-of-concept to production.
- Build and support software and product features around the inference stack in production using general-purpose languages, preferably Python.
- Experiment quickly by shaping fuzzy goals into clear specs and focused proofs of concept, running fast experiments, and shipping well-tested results.
- Own delivery end-to-end, from experiment to production optimization, maintaining focus on performance goals and drafting product requirement documents.
- Work through ambiguity, make sound decisions on tradeoffs and tooling, and steer away from unnecessary complexity.
- Take pride and ownership in work, hold yourself accountable, and expect the same from colleagues.
Requirements
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
- Hands-on experience shipping code in production with one or more general-purpose languages, such as Python or C++, with a strong preference for Python.
- Familiarity with methods for optimizing LLMs for high throughput / low latency inference.
- Comfort with modern LLM serving frameworks such as vLLM or SGLang, and with profiling and analyzing performance down to the kernel level.
- A firm grasp of how GPUs are built and how they behave.
- Clear interest and hands-on experience with large language models.
- Working knowledge of AI/ML pipelines and the full path of developing and deploying ML models.
- Strong communication skills, particularly when explaining hard technical topics to customers and teammates.
Benefits
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off, paid holidays & leave of absence programs
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend