Engineering Manager, Forward Deployed Engineering (LLM)
Remote • San Francisco • FullTime
Posted 4mo ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
LLM Engineer
About the job
Baseten is seeking an Engineering Manager (Player & Coach) to lead a team of Forward Deployed Engineers focused on building, scaling, and optimizing LLM inference workloads for their customers. This role involves both hands-on technical ownership and managerial leadership, guiding the team in designing, deploying, and managing high-performance, low-latency AI applications on Baseten’s platform. The Forward Deployed Engineering team contributes to the core Baseten codebase, influences the feature roadmap, and executes complex customer engagements. You will also collaborate with product, infrastructure, and other customer engineering teams to ensure generative AI systems deliver best-in-class performance, reliability, and cost efficiency.
Responsibilities
- Lead, mentor, and grow a team of Forward Deployed Engineers, providing guidance on technical direction, project execution, and professional development.
- Set clear goals and ensure timely, high-quality delivery across multiple customer-facing projects involving LLM deployment and inference optimization.
- Collaborate with leadership to align team priorities with company and customer goals, balancing short-term delivery, customer priorities, and long-term technical initiatives.
- Act as a player-coach, driving strategic product initiatives and customer engagements while being hands-on when needed.
- Develop and maintain software systems and product features using production-level programming languages, with a preference for Python.
- Drive customer impact by designing, implementing, and deploying Baseten solutions end-to-end, from problem framing to monitoring.
- Turn vague objectives into clear specs and well-defined Proofs of Concept (PoCs) for rapid, well-tested service delivery.
- Optimize and enhance AI/ML projects, contributing to the continuous improvement of the technical stack.
- Own products and customer projects end-to-end, functioning as an engineer, project manager, and product manager with a focus on user empathy and execution.
Requirements
- Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, or related field.
- 4+ years of professional software engineering experience.
- 1+ year in a leadership or mentorship capacity.
- Strong programming skills in Python, with production experience in building or optimizing ML inference systems.
- Proven experience with LLMs, inference optimization, or serving frameworks (e.g., vLLM, TensorRT, Triton, Hugging Face, Ray Serve).
- Familiarity with observability, profiling, and cost/performance tradeoffs in production ML systems.
- Excellent communication and collaboration skills, able to lead cross-functional efforts in ambiguous, fast-paced environments.
Benefits
- Competitive compensation, including meaningful equity
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company-wide Winter Break
- Paid parental leave
- Fertility and family-building stipend
- Company-facilitated 401(k)
- Exposure to a variety of ML startups, offering learning and networking opportunities