Safeguards Enforcement Analyst, Integrity & Authenticity
Remote • Remote-Friendly, United States; San Francisco, CA | New York City, NY | Washington, DC
Posted 16d ago
About the job
Anthropic is seeking a Safeguards Analyst focused on Integrity & Authenticity to build and execute enforcement workflows for their AI products. This role will concentrate on detecting and mitigating misuse of AI systems for coordinated inauthentic behavior, election manipulation, and targeting/surveillance. The work involves addressing AI-enabled influence operations, disinformation campaigns, electoral interference, and the use of AI for stalking and profiling. The analyst will play a key role in shaping policy enforcement to ensure safe and beneficial AI interactions. Please note, this role may involve exposure to sensitive content and require weekend/holiday availability, especially during major electoral events.
Responsibilities
- Design and architect scalable, accurate automated enforcement systems and review workflows.
- Partner with Engineering and Data Science to optimize detection models and enforcement systems.
- Review flagged content to improve enforcement and policies.
- Enforce usage policies against AI-enabled influence operations, inauthentic behavior, election interference, and surveillance.
- Provide feedback to the Safeguards policy design team based on enforcement scenarios.
- Stay updated on AI policy enforcement best practices, threat tactics, and relevant regulations.
Requirements
- Experience in trust & safety, policy enforcement, or threat intelligence, focusing on influence operations, disinformation, inauthentic behavior, election integrity, or privacy/surveillance harms.
- Experience establishing and scaling policy enforcement or content review workflows.
- Proficiency in SQL or other data analysis tools for large datasets.
- Experience identifying emerging risks and threat actors, and communicating findings to diverse stakeholders.
- Experience with generative AI products, including prompt writing for content review.
- Understanding of challenges in implementing product policies at scale, particularly in content moderation.