Red Team Engineer, Safeguards
Remote • Remote-Friendly (Travel Required) | San Francisco, CA
Posted 16d ago
Remote Work Policy
Fully remote
Categories
Applied AI Engineer
About the job
Anthropic's Safeguards team is looking for a Red Team Engineer to ensure the safety of deployed AI systems and products. This role involves taking an adversarial approach to identify vulnerabilities before they can be exploited, covering technical infrastructure and emergent risks from advanced AI capabilities. While drawing on traditional security practices, the focus is on AI-specific safety implications and novel abuse scenarios. You will investigate potential misuse, from account manipulation to exploiting product features, and simulate sophisticated threat actors using multiple attack vectors.
Responsibilities
- Conduct adversarial testing across product surfaces with creative attack scenarios.
- Research and implement novel testing approaches for emerging capabilities like agent systems and tool use.
- Design and execute full kill chain attacks emulating real-world threat actors.
- Build and maintain systematic testing methodologies for comprehensive system evaluation.
- Develop automated testing frameworks for continuous assessment at scale.
- Collaborate with Product, Engineering, and Policy teams to implement improvements.
- Help establish metrics for measuring detection effectiveness of novel abuse.
Requirements
- Experience in penetration testing, red teaming, or application security.
- Experience in model jailbreaking and testing large-scale agentic workflows for prompt injection vectors.
- Strong technical skills in web application security and security testing tools (e.g., Burp Suite, Metasploit, custom scripting).
- Experience building custom automation, including LLM-specific testing frameworks.
- Proven track record of discovering novel attack vectors and chaining vulnerabilities.
- Public body of work (CVEs, blog posts, bug bounty reports).
- Strong written and verbal communication skills for explaining technical concepts to diverse audiences.
Benefits
- Annual compensation range: $320,000 - $405,000 USD
- Visa sponsorship available
- Encouragement to apply even if not all qualifications are met
- Support for individuals from underrepresented groups