Research Product Manager, Model Behaviors
San Francisco, CA | New York City, NY
Posted 16d ago
Remote Work Policy
On-site
Categories
AI Research Engineer
About the job
Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. As a Product Manager for Model Behaviors, you will partner with the Alignment Finetuning team to define and shape Claude's character, behaviors, and reinforcement signals. This role directly influences how millions of people experience AI by systematically identifying high-priority behavioral improvements, coordinating across Research, Product, and Safeguards teams, and accelerating the ability to ship well-aligned models. The ideal candidate possesses deep user empathy and the judgment to navigate nuanced behavior questions.
Responsibilities
- Define behavioral defaults and steerability constraints.
- Develop and maintain taxonomies of model behaviors across capabilities.
- Identify, triage, and prioritize behavior issues and opportunities, coordinating input from Users, Research, Product, and Safeguards teams.
- Amplify alignment research breakthroughs, translating them into product, process, and model improvements.
- Deeply understand user interaction patterns to identify behavior improvements that make Claude more helpful and safe.
- Contribute to evaluations that measure alignment progress.
- Identify and scale initiatives and tools that help researchers ship alignment improvements faster.
Requirements
- Deep passion and curiosity for AI and LLMs; use AI regularly.
- 5+ years in product management leading scaled conversational AI products.
- First-principles thinker with the ability to navigate and execute amidst ambiguity.
- Track record of delivering products and features to end-users (consumer or end-user b2b focus).
- Strong user empathy and the ability to synthesize vague or contradictory feedback into actionable priorities.
- Strong judgment and model taste, with the ability to make tradeoffs when there is no clear right answer.
- Strong grasp of ML concepts and willingness to go deep on technical solutions.
- Intellectual curiosity without ego; comfortable asking questions and learning independently.
- Creative thinking about the risks and benefits of new technologies.
- Creative, hacker spirit and enjoyment of solving puzzles.
Benefits
- Annual compensation range: $385,000 - $460,000 USD
- Visa sponsorship available for some roles.
- Encouragement to apply even if not all qualifications are met.
- Commitment to diversity and inclusion.