[Expression of Interest] Research Engineer / Scientist, Alignment - London

London, UK

Posted 17d ago

Remote Work Policy

On-site

Categories

AI Research Engineer

About the job

Anthropic is building reliable, interpretable, and steerable AI systems to ensure AI is safe and beneficial for society. As a Research Engineer on the Alignment Science team in London, you will design and execute machine learning experiments to understand and steer the behavior of advanced AI systems. You will focus on AI safety, particularly risks from future human-level AI systems, collaborating with teams like Interpretability and Frontier Red Team. The role involves exploratory research in areas such as AI Control and Alignment Stress-testing, aiming to make AI helpful, honest, and harmless.

Responsibilities

  • Design and run machine learning experiments to understand and steer AI behavior.
  • Contribute to exploratory experimental research on AI safety, focusing on risks from powerful future systems.
  • Collaborate with other teams including Interpretability, Fine-Tuning, and the Frontier Red Team.
  • Develop methods to ensure advanced AI systems remain safe and harmless in adversarial scenarios.
  • Create model organisms of misalignment to improve empirical understanding of alignment failures.
  • Test the robustness of safety techniques by training language models to subvert them.
  • Run multi-agent reinforcement learning experiments, such as AI Debate.
  • Build tooling to evaluate the effectiveness of LLM-generated jailbreaks.
  • Write scripts and prompts to generate evaluation questions for model reasoning abilities.
  • Contribute ideas, figures, and writing to research papers, blog posts, and talks.
  • Run experiments that inform key AI safety efforts, like the Responsible Scaling Policy.

Requirements

  • Significant software, ML, or research engineering experience.
  • Experience contributing to empirical AI research projects.
  • Familiarity with technical AI safety research.
  • Ability to pick up slack and work collaboratively in fast-moving projects.
  • Interest in the impacts of AI.
  • Experience authoring research papers in machine learning, NLP, or AI safety (preferred).
  • Experience with LLMs (preferred).
  • Experience with reinforcement learning (preferred).
  • Experience with Kubernetes clusters and complex shared codebases (preferred).
  • All interviews conducted in Python.

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.