Research Engineer (Senior Staff+), Safeguards Labs
San Francisco, CA | New York City, NY
Posted 16d ago
About the job
Safeguards Labs is a new team at Anthropic focused on AI safety, operating at the intersection of research and engineering. The team's mission is to investigate novel safety methods to protect users and the AI system itself. We prototype new approaches to safe models, usage safeguards, and production safety, pressure-testing ideas through offline analysis and subsets of traffic before they are implemented in production systems. This role involves defining and executing the Labs research agenda, scoping projects, running experiments end-to-end, and deciding on the viability of research ideas. The team is structured with a high ratio of researchers to software engineers, offering substantial autonomy and influence over the team's direction.
Responsibilities
- Lead and contribute to research projects on detecting Claude misuse, identifying malicious actors, and strengthening model safeguards.
- Design and execute offline analyses of model usage data to identify abuse patterns, build classifiers, and evaluate detection system effectiveness.
- Develop and iterate on prototypes for real-time safeguards, collaborating with engineers on technology transfer.
- Contribute to research on detecting abusive behavior in chat and agentive workflows, and training models to avoid dangerous responses without over-refusal.
- Build evaluation methodologies to measure the effectiveness of safeguards, particularly in agentic settings.
- Clearly document and communicate research findings to inform decisions across Trust & Safety, research, and product teams.
Requirements
- Proven track record of independently driving research projects from ambiguous problems to concrete results, preferably in AI, ML, security, or integrity.
- Comfortable scoping own work and adapting between research, engineering, and analysis.
- Working familiarity with large language model operations (sampling, prompting, training).
- Proficiency in Python and experience working with large datasets.
- Commitment to AI safety and reducing real-world harm.
- Experience building and training machine learning models, including classifiers for abuse, fraud, integrity, or security.
- Knowledge of evaluation methodologies for language models and experience designing evaluations.
- Experience with agentic environments and evaluating model behavior within them.
- Background in trust and safety, integrity, fraud detection, threat intelligence, or adversarial ML.
- Experience with red teaming, jailbreak research, or interpretability methods.
Benefits
- Annual compensation range: $350,000 - $850,000 USD
- Visa sponsorship available
- Hybrid work policy (at least 25% in office)