Research Engineer / Scientist, Alignment
San Francisco, CA
Posted 17d ago
Job Location
San Francisco, CA
Tech Stack
Remote Work Policy
On-site
Categories
AI Research Engineer
About the job
Anthropic is seeking a Research Engineer/Scientist for its Alignment Science team. This role involves designing and executing machine learning experiments to understand and steer the behavior of advanced AI systems, with a focus on AI safety and potential risks from future human-level AI. You will collaborate with other teams on exploratory research, contributing to Anthropic's mission of creating reliable, interpretable, and steerable AI systems that are helpful, honest, and harmless.
Responsibilities
- Design and run machine learning experiments to understand and steer AI behavior.
- Conduct exploratory experimental research on AI safety, focusing on risks from powerful future systems.
- Collaborate with teams like Interpretability, Fine-Tuning, and Frontier Red Team.
- Develop techniques for scalable oversight to keep highly capable models helpful and honest.
- Create methods to ensure advanced AI systems remain safe and harmless in adversarial scenarios.
- Develop alignment stress-testing methodologies and model organisms of misalignment.
- Build and align systems to accelerate alignment research.
- Assess and document emerging properties of models through pre-deployment alignment and welfare assessments.
- Develop robust defenses against adversarial attacks and comprehensive evaluation frameworks for model safety.
- Investigate and address model welfare, moral status, and related questions.
- Test the robustness of safety techniques by training models to subvert them.
- Run multi-agent reinforcement learning experiments, such as AI Debate.
- Build tooling to evaluate the effectiveness of LLM-generated jailbreaks.
- Write scripts and prompts to generate evaluation questions for model reasoning in safety contexts.
- Contribute ideas, figures, and writing to research papers, blog posts, and talks.
- Run experiments that support key AI safety efforts, including the design and implementation of safety techniques.
Requirements
- Ability to build and run elegant and thorough machine learning experiments.
- Interest in making AI helpful, honest, and harmless.
- Interest in challenges related to human-level AI capabilities.
- Proficiency in both scientific and engineering aspects of AI research.
- Interviews conducted in Python.
- Preference for candidates based in the Bay Area.