Engineering Manager, Evals
San Francisco • FullTime
Posted 3mo ago
Job Location
San Francisco
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
As an Engineering Manager on the Evals team at Cursor, you will lead the group responsible for creating high-signal evaluation datasets for coding agents and building the tools engineers use to write and run them. This team also owns online evaluation systems that track agent quality in production, and the close integration between online and offline evaluations. The evaluation systems that this team builds are critical in the development of our coding models and the quality of our Cursor agents. Your impact will compound across every Cursor product and every Cursor model by making quality measurable, comparable, and easy to improve.
Responsibilities
- Set the eval roadmap end-to-end, defining what we measure, why it matters, and how signals turn into shipping and training decisions.
- Lead and grow a high-impact team of engineers and researchers building eval datasets and developer-friendly tools to write and run evals.
- Guide the next generation of CursorBench to reflect real developer workflows and expand it with new evals that measure other properties developers value.
- Define crisp online quality signals and turn regressions into robust guardrails.
- Integrate evals into decision-making cadence for launches, deploys, and model training loops.
Requirements
- Experience leading engineering teams shipping production systems with strong people leadership and coaching skills.
- Ability to align research, product, data, and infrastructure on what “good” means and translate that into durable metrics, processes, and release/training rituals.
- Good taste and strong opinions on model and agent behaviors, with up-to-date knowledge of emerging research and industry trends.
- Strong data acumen and ability to collaborate effectively with data scientists and researchers.
- Experience building and operating evaluation or measurement systems (e.g., AI evals, experimentation platforms, ranking/relevance, search quality, or reliability instrumentation).