Researcher, Agent Safety, Oversight and System Mitigations
San Francisco • FullTime
Posted 10d ago
About the job
This role focuses on oversight and system-level mitigations that enable increasingly capable agents to operate safely and autonomously in real environments. We prioritize building oversight systems that are used in practice today, both internally and externally. We also study longer-term questions about how increasingly capable agent systems can be supervised, constrained, and corrected. We’re looking for a safety and security-minded researcher or engineer who can reason rigorously about security boundaries and agent behavior, then build and test practical mitigations. A background in AI control or security is welcome but not required.
Responsibilities
- Design, build, and evaluate system-level controls for agent actions, such as agent-based review, and plan their integration into broader systems.
- Collaborate with a Codex harness engineering team to productionize AI controls.
- Conduct red-teaming of end-to-end agentic systems to assess the effectiveness of controls against various harmful outcomes.
- Improve the safety-productivity tradeoff by measuring and reducing missed harmful actions, unnecessary blocks, approval burden, and latency.
Requirements
- Strong systems or security instincts with the ability to reason concretely about isolation boundaries, permissions, attack surfaces, and failure modes in complex systems.
- Ability to translate ambiguous safety questions into concrete threat models, reproducible experiments, and practical mitigations.
- Capacity to build robust experimental infrastructure and design evaluations that distinguish effective mitigations from less reliable ones.
- Deep interest in frontier AI alignment, safety, and control.
Benefits
- Relocation assistance to new employees.