Software Engineer, Agent Evaluation and Quality
San Francisco • FullTime
Posted 5mo ago
Job Location
San Francisco
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Agent Engineer
About the job
As a Software Engineer on the Agent Quality team, you will build the measurement, evaluation, and feedback-loop infrastructure that makes the core agent reliably better over time. This role sits at the intersection of product, data, and engineering. You will instrument what matters, help define how quality is judged, build pipelines and tooling to analyze agent behavior at scale, and partner closely with research, product, and infrastructure teams to turn insights into improvements. Your impact will compound across every product built on the shared harness and across high-stakes decisions around model choice, quality, and cost.
Responsibilities
- Design and build AI evaluation systems, including curated datasets, offline replay, scorers/judges, regression alerts, and dashboards.
- Design feedback loops from real usage, collecting, cleaning, and interpreting user signals to inform model and harness changes.
- Develop analysis tooling and workflows for debugging agent behavior, including deep dives on failure modes, clustering themes, and surfacing actionable insights.
- Improve reliability and guardrails by making quality measurable and operational, defining "good/bad/degraded" sessions, alerting, and triage primitives.
Requirements
- Experience building and operating evaluation or measurement systems (e.g., AI evals, experimentation, ranking/relevance, search quality).
- Ability to translate ambiguous "quality" questions into concrete metrics, pipelines, and decisions.
- Strong data acumen and ability to collaborate effectively with data scientists and researchers.
- Taste and strong opinions on model and agent behaviors, staying up-to-date on emerging research and industry trends.
- Strong software engineering fundamentals and experience shipping production systems.