Research Engineer, Domain Scaling
San Francisco, CA | New York City, NY | Seattle, WA
Posted 16d ago
Job Location
San Francisco, CA | New York City, NY | Seattle, WA
Tech Stack
Remote Work Policy
On-site
Categories
AI Research Engineer
About the job
The Domain Scaling team aims to make Claude world-class at real-world knowledge work in domains like finance, healthcare, and legal. This role combines direct applied research with data sourcing (real-world and synthetic) to improve our models. You will own the end-to-end process of creating RL environments for new capabilities, which includes identifying high-value tasks, designing reward signals, managing vendor relationships, and measuring impact on model performance.
Responsibilities
- Own the data strategy for knowledge work verticals end-to-end, from task sourcing through RL training.
- Manage technical relationships with external data vendors, including evaluation of data quality and reward design.
- Collaborate with domain experts to design data pipelines and evaluations.
- Explore novel ways of creating RL environments for high-value tasks.
- Develop and improve QA frameworks to catch reward hacking and ensure environment quality.
- Run generalization experiments to measure how data strategy changes improve model capabilities.
- Partner with other RL research teams and product teams to translate capability goals into training environments and evaluations.
Requirements
- Experience with fine-tuning large language models for specific domains or real-world use cases.
- Experience with reinforcement learning, reward design, or training data curation for LLMs.
- Comfortable managing technical vendor relationships and iterating quickly on feedback.
- Ability to read through datasets to understand them and spot issues.
- Strong cross-functional collaboration skills.
- Passion for making AI more useful and accessible across different industries.
- Excited about a role that includes a combination of applied research and hands-on data work.
- Experience training production ML systems.
- Experience designing evaluations or benchmarks for LLMs.
- Domain expertise in a vertical where we would like to make our models more useful.
- Experience working with external vendors or technical partners.