Research Engineer – Benchmarking

San Francisco FullTime

Posted 27d ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Research Engineer

About the job

As a Research Engineer at Mercor, you will operate at the intersection of engineering and applied AI research. Your primary focus will be on owning benchmarking pipelines, evaluation systems, and failure analysis workflows that directly influence how frontier language models are trained and improved. Your work will be instrumental in defining how we measure tool use, agentic behavior, and real-world reasoning. You will be responsible for designing and executing evaluations, building rubrics and scoring mechanisms, and translating failure analysis findings into actionable improvements for post-training, RLVR, and data pipelines.

Responsibilities

  • Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning, ensuring scalability and alignment with product and research goals.
  • Build and operate end-to-end LLM evaluation systems, including scoring, dashboards, and reporting, to track model performance and facilitate comparisons.
  • Conduct systematic failure analysis on model outputs, categorizing and quantifying failure modes to inform reward design, data curation, and benchmark design.
  • Create and refine rubrics, automated evaluators, and scoring frameworks to guide training and evaluation decisions, balancing rigor with scalability.
  • Quantify data usability, quality, and impact on key benchmarks, using evaluations and failure analysis to guide data generation and curation.
  • Collaborate with AI researchers, applied AI teams, and data producers to align evaluations with training objectives and prioritize impactful benchmarks and analyses.
  • Take ownership of benchmarks, evaluations, and failure-analysis workflows in a high-iteration research setting.

Requirements

  • Strong applied research background with a focus on model evaluation, benchmarking, and/or failure analysis.
  • Strong coding skills and hands-on experience with ML models and evaluation code.
  • Solid understanding of data structures, algorithms, and backend systems.
  • Comfort with APIs, SQL/NoSQL, and cloud platforms for running and storing evaluation results.
  • Ability to reason about model behavior, experimental results, and data quality from evaluations and failure analyses.
  • Experience building or running LLM evaluations, benchmarks, or failure-analysis pipelines.
  • Experience with synthetic data generation, rubric design, or RL-style workflows that use evals for reward shaping.

Benefits

  • Bi-annual performance bonus structure
  • Generous equity grant vested over 4 years
  • Up to $15k Relocation bonus
  • $10K housing bonus (if you live within 0.5 miles of our office)
  • $1.5K monthly stipend for meals
  • Free Equinox membership
  • $200 monthly laundry reimbursement
  • $200 monthly personal wellness reimbursement
  • Health, Dental, Vision insurance

About Mercor

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.