Inference Engineering and Product Lead
Remote • San Francisco • FullTime
Posted 11h ago
Job Location
San Francisco
Tech Stack
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
LLM Engineer
About the job
Modal is building the next-generation infrastructure layer for AI, focusing on serving LLM inference with frontier performance and best-in-class elasticity. We are seeking a leader to own the direction and execution of our LLM inference platform, working closely with talented engineers and high-profile customers. This is a hands-on leadership role where you will split your time between technical contributions, product shaping, and people management. You will set the team's direction, remove obstacles, and cultivate a strong engineering culture while tackling complex challenges in distributed computing, inference serving, and performance optimization.
Responsibilities
- Recruit, hire, and grow a high-performing engineering team, providing coaching and career development.
- Set clear performance expectations and foster a culture of ownership, accountability, and customer obsession.
- Drive technical and product decisions through design reviews, code reviews, and architectural discussions.
- Lead customer engagements for novel or frontier workloads to ensure their success on Modal.
- Translate customer learnings into a roadmap for internal optimization platforms and user-facing products.
- Establish standards for reliability and product excellence, ensuring end-to-end project ownership.
- Partner with business operations and compute strategy on compute purchase strategies.
- Collaborate with Go-to-Market teams on product launches and inference opportunity win rates.
- Guide the roadmap for underlying inference infrastructure and adjacent product teams.
Requirements
- 10+ years of industry experience, with at least 3 years in a leadership role.
- Proven track record of building high-performance systems at scale.
- Strong background in cloud infrastructure.
- Deep knowledge of low-level OS foundations (Linux kernel, file systems, containers).
- Experience with LLM inference in production is a plus.
- Familiarity with concepts like inference engines, kernels, routing, KV cache management, and speculative decoding is a plus.