Backend Software Engineer (Evals)
San Francisco • FullTime
Posted 8mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
We are seeking a Backend Software Engineer with experience in ML/LLM-heavy domains to design and build an "evals" infrastructure. This role is critical for measuring the quality of OpenAI's support automation and involves building robust systems and backend services that form the foundation for knowledge creation, access, and application across OpenAI. You will collaborate closely with Data Science and Research partners to design and build evaluations at scale, focusing on creating reliable, reproducible, and extendable eval pipelines and infrastructure for continuous monitoring.
Responsibilities
- Design reliable, reproducible, and extendable eval pipelines.
- Build infrastructure for continuous eval monitoring frameworks, including regression/drift monitoring and robust golden datasets.
- Develop feedback loops to strengthen support automation.
- Design, build, and maintain backend services and APIs for intelligent automation and knowledge systems.
- Integrate and structure data across internal platforms for downstream systems and AI workflows.
- Collaborate with data, research, and engineering teams to integrate OpenAI models into high-leverage workflows.
- Own the full development lifecycle of new backend systems and internal platform capabilities.
- Build with scale and maintainability in mind, while rapidly iterating on new ideas.
Requirements
- 4+ years of backend engineering experience at product-driven companies (excluding internships).
- Proficiency in backend technologies, including Python, FastAPI, and Postgres.
- Experience designing and scaling distributed systems, APIs, or data processing pipelines.
- Experience building AI agents or applications, including designing evals and improving performance through prompting or scaffolding.
- Familiarity with evaluation methods for LLMs and experience with patterns like multi-agent workflows, tool use, or long context.
- Experience creating production evals and/or measuring performance of ML/LLM models at scale.
- A pragmatic mindset, comfortable shipping iteratively while building toward a long-term vision.