Research Scientist, APEX Benchmarks
San Francisco • FullTime
Posted 17d ago
About the job
Mercor is seeking a Research Scientist to lead the design of the next generation of APEX benchmarks and the expert-built datasets that support them. APEX benchmarks measure the real-world economic impact of frontier AI models on professional work, with current benchmarks covering areas like investment banking, law, consulting, medicine, accounting, and software engineering. This role is highly visible, sitting at the intersection of research, company strategy, and go-to-market efforts. You will be responsible for identifying future measurement areas, designing benchmarks and scoring systems, collaborating with partners to build them, and working cross-functionally to scale production. You will also serve as a credible external technical voice, engaging with labs, partners, and the research community. This is an opportunity to build new methodologies, tooling, and publishing practices in a fast-paced, ambitious environment.
Responsibilities
- Design the next generation of APEX benchmarks, determining what to measure based on frontier model performance and gaps in existing evaluations.
- Own benchmark design, including task taxonomy, difficulty calibration, contamination controls, and statistical design.
- Design expert-built datasets and grading rubrics at scale, defining task difficulty and defensible grading criteria.
- Set standards for reporting results, including confidence intervals, inter-rater agreement, and failure analysis.
- Collaborate with academic and industry partners to co-design benchmarks and drive adoption.
- Publish research findings through papers, datasets, blog posts, and conference talks.
- Represent Mercor's research externally to frontier labs, customers, and the press.
- Translate benchmark findings into clear arguments about the ROI of expert-curated data for technical reports and go-to-market materials.
- Partner with data operations, engineering, product, and strategy teams to move benchmarks from design to production.
- Track LLM evaluation literature and integrate relevant advancements into Mercor's benchmark development processes.
Requirements
- Strong applied or academic research background in LLM evaluation, benchmarking, NLP, or a related field with a track record of rigorous experimental design.
- Ability to identify informative measurements for frontier models based on their behavior.
- Proficiency in statistical rigor, including reasoning about sampling, variance, contamination, and grader reliability.
- Strong coding skills for building evaluation harnesses, running experiments, and analyzing results.
- Exceptional communication skills, capable of presenting complex technical findings clearly to both technical and non-technical audiences.
- Comfort operating in fast-moving, cross-functional environments with undefined problem spaces.
- Genuine curiosity about go-to-market strategy, startup dynamics, and the business of AI data.
- Willingness to work in-person five days a week in the San Francisco office.