Research Engineer, Post-Training
$231k - $340k • Remote • San Francisco • FullTime
Posted 1mo ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Research Engineer
About the job
Harvey is transforming legal and professional services by combining agentic AI, an enterprise-grade platform, and deep domain expertise. We are seeking a Research Engineer focused on post-training to scale the process of turning expert feedback and agent traces into significantly improved models. This role involves defining and running model training experiments, interpreting results, and collaborating with internal and external partners to enhance data, environments, graders, and training methodologies. The ideal candidate is a self-manager with extensive hands-on experience training open-weight models and the engineering depth to execute and debug experiments efficiently.
Responsibilities
- Drive post-training experiments to improve agent performance while balancing cost, latency, security, and governance.
- Optimize agent harnesses, including domain-specific skills, tools, subagents, retrieval strategies, and validation loops for complex legal work.
- Design and develop reliable grading and reward systems for evaluation and iteration in high-stakes legal contexts.
- Analyze agent behavior to identify patterns correlating with successful work product and translate findings into training data, evaluations, or harness improvements.
- Collaborate with researchers and external partners to define experiments, evaluate methodologies, review results, and drive concrete model improvements.
Requirements
- Hands-on experience with post-training or model-training techniques like SFT, preference optimization, RLHF/RLAIF, reward modeling, distillation, or adapting open-weight models.
- Strong judgment regarding model behavior, including trace analysis, output inspection, failure mode identification, and metric relevance.
- Proficiency in Python and research-engineering, with the ability to write clean code, debug experiments, and build reliable research systems.
- Ability to self-manage ambiguous applied research projects and communicate effectively with diverse teams and external partners.