Safeguards Enforcement Analyst, Conventional Weapons
San Francisco, CA | New York City, NY | Washington, DC
Posted 2d ago
About the job
Anthropic is building reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. As a Safeguards Enforcement Analyst focused on Conventional Weapons, you will play a crucial role in detecting and mitigating attempts to misuse Anthropic's AI systems for real-world harm, specifically involving conventional weapons and dangerous technology. This position involves building and executing operational workflows to assess model behavior, making enforcement decisions, and developing evaluations across technically complex policy areas. Please note that this role may involve exposure to explicit content of a violent, graphic, hateful, or psychologically disturbing nature.
Responsibilities
- Design and architect automated enforcement systems and review workflows for accuracy and scalability.
- Develop and maintain evaluations to measure model performance on policy areas, identify regressions, and inform improvements.
- Collaborate with Engineering and Data Science to optimize detection and automated enforcement systems for policy violations.
- Review flagged content to make enforcement decisions, identify policy gaps, and address novel misuse attempts.
- Provide structured feedback to the Safeguards policy design team on policy gaps and enforcement ambiguities.
- Develop and maintain enforcement guidelines and reviewer documentation for consistent enforcement.
- Stay updated on emerging weapons trends, regulatory changes, and AI policy enforcement best practices.
- Identify and escalate emerging misuse patterns, novel attack vectors, and signs of coordinated violent activity.
Requirements
- Deep, applied expertise in weapons systems and ability to translate technical evidence for enforcement decisions.
- Experience in policy enforcement, threat intelligence, counterterrorism, government, or a related field with exposure to harmful content or dangerous technology.
- Experience establishing and scaling policy enforcement or content review workflows.
- Proficiency in SQL and/or other data analysis tools for insights and workflow monitoring.
- Experience identifying emerging risks and threat actors, and communicating findings to diverse stakeholders.
- Experience working with generative AI products, including prompt writing for content review and enforcement.
- Understanding of challenges in implementing product policies at scale, particularly in content moderation.