Forward Deployed Engineer (Training)
Remote • San Francisco • FullTime
Posted 26d ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Baseten is seeking a Forward Deployed Engineer (Training) to work directly with leading AI companies, taking ownership of their technical outcomes on the Baseten platform. This role involves tackling complex challenges in serving and improving AI models at scale, spanning the entire model lifecycle from inference to post-training and the systems that connect them. You will act as a technical advisor, guiding customers from initial problem framing through to production deployment, and ensuring the quality and performance of their AI workloads through rigorous evaluation and optimization.
Responsibilities
- Serve as the de facto CTO for customer accounts on Baseten, ensuring their workloads are designed, run, and scaled effectively.
- Guide customer objectives from vague ideas to shipped products by defining problems, setting success criteria, building proofs-of-concept, and managing the transition to production.
- Design and implement evaluations and benchmarks to identify quality or performance issues, and then resolve these issues through inference optimization, post-training improvements, or eval rework.
- Respond to critical failures, including triage, owning the fix, or coordinating with the appropriate teams, ensuring accountability until resolution.
- Develop internal systems, tooling, and automation to improve the efficiency of evaluations and deployment infrastructure, and create self-serve resources.
- Contribute to the Baseten product roadmap by channeling customer needs into feature development and bug fixes within the Baseten codebase.
- Manage multiple customer engagements simultaneously, prioritizing work, coordinating resources, and maintaining alignment with customers and internal stakeholders on status and risks.
Requirements
- Minimum 1-2 years of software engineering experience, with a track record of shipping and maintaining code in large production systems, ideally with full-stack exposure.
- Proven ability to debug complex production issues using logs, metrics, and traces to identify root causes in unfamiliar systems.
- Confidence in owning ambiguous technical problems, including triaging, making decisions with incomplete information, and knowing when to escalate.
- Motivation to work directly with customers, understand their challenges, and influence product development.
- Clear communication skills for explaining complex technical topics to both engineers and leadership.
- Genuine curiosity about AI inference and training, with a drive to become an expert in AI infrastructure.
- Willingness to be available for customer support outside regular hours and participate in an on-call rotation.
- Enthusiasm for solving problems for major AI companies running mission-critical workloads.
Benefits
- Competitive compensation package including meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employees and dependents.
- Flexible Paid Time Off (PTO) policy.
- Company-wide Winter Break (offices closed from Christmas Eve to New Year's Day).
- Paid parental leave.
- Fertility and family-building stipend.
- Company-facilitated 401(k).
- Exposure to diverse ML startups for learning and networking opportunities.