Engineering Manager, AI Platform
$215k - $260k • San Francisco, CA - US • FullTime
Posted 1mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
As an Engineering Manager on the Managed AI team, you will lead and scale a team of engineers building our next-generation platform for the full lifecycle of Large Language Models (LLMs). You will guide the team through the design and implementation of highly scalable, fault-tolerant infrastructure, combining technical expertise with strong people leadership. This role is central to a fast-growing, strategically important organization where you will shape the engineering roadmap and drive the execution of key projects. Success requires close partnership with product, business, and platform stakeholders to deliver a performant and reliable platform that powers AI for customers globally.
Responsibilities
- Lead, mentor, and grow a team of software engineers.
- Partner with leadership to define and execute the AI roadmap, setting clear goals and driving accountability.
- Cultivate a high-performance, collaborative engineering culture.
- Oversee the architecture and development of core AI services like fault-tolerant task queues and model management systems.
- Ensure delivery of scalable systems capable of handling millions of API requests per second.
- Deliver an AI platform that can handle a large variety of loads, from training to agentic execution infrastructure.
- Work cross-functionally with Product, Infrastructure, and GTM stakeholders.
- Represent Engineering in strategic discussions to influence AI platform growth and customer adoption.
- Promote knowledge sharing, technical mentorship, and the evolution of engineering processes.
Requirements
- 3+ years managing/leading high-performing engineering teams.
- Ability to lead teams through ambiguity and align on complex technical goals.
- Proven success hiring, developing, and retaining talent.
- Hands-on experience with distributed and concurrent systems or AI infrastructure.
- Deep knowledge of cloud-native environments, container orchestration, and SOAs.
- Familiarity with CPU & GPU performance, inference frameworks, or LLM systems is a strong plus.
- Comfortable owning deliverables from design through production.
- Strong collaboration skills, prioritizing clarity, context, and customer impact.
- Experience in fast-paced startup or growth-stage environments.
- Proficiency in Python/GoLang/Rust (preferred).
- Experience with Kubernetes, gRPC, and observability stacks (preferred).
- Familiarity with open-source AI ecosystems (e.g., vLLM, Hugging Face, Triton) (preferred).
- Growth-minded leader who leads and empowers others.
- Excellent communicator and relationship builder.
- Passionate about building world-class AI infrastructure and teams.
Benefits
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off, paid holidays & leave of absence programs
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance & emergency assistance
- Daily meals allowance
- Additional perks & programs specific to location