Research Engineer - Environments, Data and Post-Training
San Francisco • FullTime
Posted 5mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Research Engineer
About the job
Mercor is a leading AI data company that organizes human intelligence to power the AI economy. As a Research Engineer, you will work at the intersection of engineering and applied AI research, contributing directly to post-training, RLVR, synthetic data generation, and large-scale evaluation workflows that impact frontier language models. Your work will be instrumental in training LLMs for tool use, agentic behavior, and real-world reasoning in production environments. You will shape rewards, run experiments, build scalable systems, and design/evaluate datasets and augmentation pipelines to enhance model performance and push the boundaries of LLM learning.
Responsibilities
- Work on post-training and RLVR pipelines to understand how datasets, rewards, and training strategies impact model performance.
- Design and run reward-shaping experiments and algorithmic improvements (e.g., GRPO, DAPO) to improve LLM tool-use, agentic behavior, and real-world reasoning.
- Quantify data usability, quality, and performance uplift on key benchmarks.
- Build and maintain data generation and augmentation pipelines that scale with training needs.
- Create and refine rubrics, evaluators, and scoring frameworks that guide training and evaluation decisions.
- Build and operate LLM evaluation systems, benchmarks, and metrics at scale.
- Collaborate closely with AI researchers, applied AI teams, and experts producing training data.
- Operate in a fast-paced, experimental research environment with rapid iteration cycles and high ownership.
Requirements
- Strong applied research background, with a focus on post-training and/or model evaluation.
- Strong coding proficiency and hands-on experience working with machine learning models.
- Strong understanding of data structures, algorithms, backend systems, and core engineering fundamentals.
- Familiarity with APIs, SQL/NoSQL databases, and cloud platforms.
- Ability to reason deeply about model behavior, experimental results, and data quality.
- Excitement to work in person in San Francisco, five days a week (with optional remote Saturdays), and thrive in a high-intensity, high-ownership environment.
- Real-world post-training team experience in industry (highest priority).
- Publications at top-tier conferences (NeurIPS, ICML, ACL).
- Experience training models or evaluating model performance.
- Experience in synthetic data generation, LLM evaluations, or RL-style workflows.
Benefits
- Bi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- $200 monthly laundry reimbursement
- $200 monthly personal wellness reimbursement
- Health, Dental, Vision insurance