Forward Deployed Engineer
Remote • San Francisco • FullTime
Posted 4h ago
Job Location
San Francisco
Tech Stack
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Agent Engineer
About the job
You will embed directly with customer teams as the engineer on the ground, owning an outcome and shipping a production-grade AI system. You will become the critical technical feedback loop from the field back to product engineering. The work involves taking messy business workflows and putting production agents on them, picking the use case, standing up the architecture, wiring it into the customer’s auth, data, and tools, and hardening it with evals until the team actually uses it. You will also live in a customer’s repo and delivery system to ship an agent-powered path for a migration, a PR review loop, or an incident-to-fix flow. You will measure what matters, leave the customer able to operate the system, and bring the pattern back as another use case.
Responsibilities
- Lead customer discovery, orienting in their codebases and workflows to find bottlenecks and define success metrics.
- Design and ship production agents, AI applications, and coding-agent workflows on Grok and other models.
- Get a first version live in days, then harden it over weeks with rollout, monitoring, and iteration.
- Own production quality, including tracing, evals, debugging model or delivery failures, and latency/cost tradeoffs.
- Build and integrate systems around the model (tools, MCP servers, retrieval, rules, skills, CI gates, evals) into customer systems.
- Measure outcomes like revenue, cost, hours saved, error rate, cycle time, and escaped defects.
- Work directly with leaders, going deep in their systems while communicating tradeoffs and results.
- Ensure the customer team can run what you built and translate learnings into reusable patterns and product improvements.
Requirements
- 5+ years of experience in software engineering, machine learning engineering, or data science.
- Proficiency in writing and reviewing production code (Python, JavaScript/TypeScript; other languages welcome).
- Experience owning a customer or operator outcome and turning fuzzy problems into scoped, shipped systems.
- Experience building and owning AI-native workflows or agents in production, including debugging real production failures.
- Experience handling production reliability: metrics, alerts, safe rollouts, incident response.
- Ability to build end-to-end systems: frontend, backend, infra, and prompt iteration.
- Thrive in ambiguity and build without a complete spec.
- Ability to communicate effectively with engineers and VPs.