Research Engineer – Benchmarking
San Francisco • FullTime
Posted 27d ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Research Engineer
About the job
As a Research Engineer at Mercor, you will operate at the intersection of engineering and applied AI research. Your primary focus will be on owning benchmarking pipelines, evaluation systems, and failure analysis workflows that directly influence how frontier language models are trained and improved. Your work will be instrumental in defining how we measure tool use, agentic behavior, and real-world reasoning. You will be responsible for designing and executing evaluations, building rubrics and scoring mechanisms, and translating failure analysis findings into actionable improvements for post-training, RLVR, and data pipelines.
Responsibilities
- Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning, ensuring scalability and alignment with product and research goals.
- Build and operate end-to-end LLM evaluation systems, including scoring, dashboards, and reporting, to track model performance and facilitate comparisons.
- Conduct systematic failure analysis on model outputs, categorizing and quantifying failure modes to inform reward design, data curation, and benchmark design.
- Create and refine rubrics, automated evaluators, and scoring frameworks to guide training and evaluation decisions, balancing rigor with scalability.
- Quantify data usability, quality, and impact on key benchmarks, using evaluations and failure analysis to guide data generation and curation.
- Collaborate with AI researchers, applied AI teams, and data producers to align evaluations with training objectives and prioritize impactful benchmarks and analyses.
- Take ownership of benchmarks, evaluations, and failure-analysis workflows in a high-iteration research setting.
Requirements
- Strong applied research background with a focus on model evaluation, benchmarking, and/or failure analysis.
- Strong coding skills and hands-on experience with ML models and evaluation code.
- Solid understanding of data structures, algorithms, and backend systems.
- Comfort with APIs, SQL/NoSQL, and cloud platforms for running and storing evaluation results.
- Ability to reason about model behavior, experimental results, and data quality from evaluations and failure analyses.
- Experience building or running LLM evaluations, benchmarks, or failure-analysis pipelines.
- Experience with synthetic data generation, rubric design, or RL-style workflows that use evals for reward shaping.
Benefits
- Bi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- $200 monthly laundry reimbursement
- $200 monthly personal wellness reimbursement
- Health, Dental, Vision insurance