Safeguards Enforcement Analyst, Violence & Extremism

Remote Remote-Friendly, United States; San Francisco, CA | New York City, NY | Washington, DC

Posted 16d ago

Job Location

Remote-Friendly, United States; San Francisco, CA | New York City, NY | Washington, DC

Tech Stack

Remote Work Policy

Fully remote

Categories

Applied AI Engineer

About the job

Anthropic is seeking a Safeguards Enforcement Analyst focused on Violence & Extremism to build and execute operational workflows for assessing AI model behavior, driving enforcement decisions, and developing evaluations across a range of policy areas. This role involves detecting and mitigating misuse of AI systems that could lead to real-world harm, including issues related to weapons, critical infrastructure, violent extremism, and threats of violence. Candidates should be aware that this position may involve exposure to explicit and disturbing content.

Responsibilities

  • Design and architect scalable automated enforcement systems and review workflows.
  • Develop and maintain evaluations to measure model performance and identify regressions.
  • Partner with Engineering and Data Science to optimize detection and enforcement systems.
  • Review flagged content to make enforcement decisions and identify policy gaps, focusing on novel misuse attempts and emerging extremist tactics.
  • Provide structured feedback to the Safeguards policy design team on policy gaps and ambiguities.
  • Develop and maintain enforcement guidelines and reviewer documentation for consistent enforcement.
  • Stay updated on emerging threats, extremist movements, regulatory changes, and AI policy enforcement best practices.
  • Identify and escalate emerging misuse patterns, novel attack vectors, and signs of coordinated violent extremist activity.

Requirements

  • Experience in policy enforcement, threat intelligence, counterterrorism, government, or a related field with exposure to harmful content or violent extremism.
  • Experience establishing and scaling policy enforcement or content review workflows.
  • Proficiency in SQL or other data analysis tools for large datasets.
  • Experience identifying emerging risks and communicating findings to diverse stakeholders.
  • Experience working with generative AI products, including prompt writing for content review.
  • Understanding of challenges in implementing product policies at scale, particularly in content moderation.

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.