Safeguards Enforcement Analyst, Ban Evasion & Recidivism

Remote Remote-Friendly, United States; San Francisco, CA | New York City, NY | Washington, DC

Posted 16d ago

Job Location

Remote-Friendly, United States; San Francisco, CA | New York City, NY | Washington, DC

Tech Stack

Remote Work Policy

Fully remote

Categories

Applied AI Engineer

About the job

As a Safeguards Enforcement Analyst on the account abuse team, you will build and execute enforcement workflows to ensure product safety, with a primary focus on detecting and mitigating potential harm. Your initial responsibilities will center on recidivism, ensuring that bans are effective and that banned actors cannot easily re-register. This role involves identifying returning banned actors, linking accounts across different identities, and closing the most critical re-registration pathways. The scope includes high-stakes areas like preventing evasion of child-safety enforcement bans, where the consequences of failure are severe. This position has the potential to expand into broader enforcement areas over time, contributing to Anthropic's mission of creating safe and beneficial AI systems.

Responsibilities

  • Investigate ban evasion clusters from initial signals to full actor networks.
  • Translate individual findings into systemic controls and detection proposals.
  • Operationalize re-registration controls for high-severity ban populations.
  • Collaborate with Engineering and Data Science on account-linking signals.
  • Develop a framework to measure recidivism and the effectiveness of controls.
  • Create playbooks for contractor-supported evasion review with quality assurance.
  • Stay updated on AI policy enforcement best practices to inform workflows.

Requirements

  • Experience investigating ban evasion, multi-accounting, or repeat fraud actors.
  • Fluency in SQL and ability to build analyses on large datasets.
  • Experience with fraud or identity-linking signals and understanding of precision/recall tradeoffs.
  • Rigor regarding evidence standards, especially for high-cost false positives.
  • Proven ability to turn one-off investigations into repeatable detection logic and policy.
  • Strong written communication skills for briefs and recommendations.
  • Excellent judgment and collaboration skills in a fast-paced environment.

Benefits

  • Annual compensation range: $245,000 - $285,000 USD
  • Visa sponsorship available
  • Hybrid work policy (at least 25% in office)

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.