Research Engineer / Scientist, Alignment

San Francisco, CA

Posted 17d ago

Remote Work Policy

On-site

Categories

AI Research Engineer

About the job

Anthropic is seeking a Research Engineer/Scientist for its Alignment Science team. This role involves designing and executing machine learning experiments to understand and steer the behavior of advanced AI systems, with a focus on AI safety and potential risks from future human-level AI. You will collaborate with other teams on exploratory research, contributing to Anthropic's mission of creating reliable, interpretable, and steerable AI systems that are helpful, honest, and harmless.

Responsibilities

  • Design and run machine learning experiments to understand and steer AI behavior.
  • Conduct exploratory experimental research on AI safety, focusing on risks from powerful future systems.
  • Collaborate with teams like Interpretability, Fine-Tuning, and Frontier Red Team.
  • Develop techniques for scalable oversight to keep highly capable models helpful and honest.
  • Create methods to ensure advanced AI systems remain safe and harmless in adversarial scenarios.
  • Develop alignment stress-testing methodologies and model organisms of misalignment.
  • Build and align systems to accelerate alignment research.
  • Assess and document emerging properties of models through pre-deployment alignment and welfare assessments.
  • Develop robust defenses against adversarial attacks and comprehensive evaluation frameworks for model safety.
  • Investigate and address model welfare, moral status, and related questions.
  • Test the robustness of safety techniques by training models to subvert them.
  • Run multi-agent reinforcement learning experiments, such as AI Debate.
  • Build tooling to evaluate the effectiveness of LLM-generated jailbreaks.
  • Write scripts and prompts to generate evaluation questions for model reasoning in safety contexts.
  • Contribute ideas, figures, and writing to research papers, blog posts, and talks.
  • Run experiments that support key AI safety efforts, including the design and implementation of safety techniques.

Requirements

  • Ability to build and run elegant and thorough machine learning experiments.
  • Interest in making AI helpful, honest, and harmless.
  • Interest in challenges related to human-level AI capabilities.
  • Proficiency in both scientific and engineering aspects of AI research.
  • Interviews conducted in Python.
  • Preference for candidates based in the Bay Area.

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.