Researcher, Frontier Risk Mitigations
San Francisco • FullTime
Posted 1mo ago
About the job
We are seeking exceptional researchers to push the frontier of safety mitigations for increasingly capable AI models. You will help derisk frontier models by developing novel safety mitigations and applying new techniques from domains like interpretability, control, and alignment. This role is critical in defining the future of safe AI systems at OpenAI and significantly impacting our mission to build and deploy safe AGI. The position requires strong technical depth and close cross-functional collaboration to ensure safety mitigations are enforceable, scalable, and effective, partnering with experts across misalignment, cybersecurity, and biology to develop a comprehensive safety stack.
Responsibilities
- Identify emerging AI safety risks and develop new methodologies for exploring and mitigating their impact.
- Build and refine evaluations to assess the extent of AI safety risks, potentially collaborating with domain experts.
- Set research directions and strategies to enhance the safety, alignment, and robustness of AI systems.
- Contribute to the development of AI safety best practices for OpenAI and the broader industry.
- Evaluate and design effective red-teaming pipelines to test the robustness of safety systems and identify areas for improvement.
Requirements
- 2+ years of experience in AI safety, particularly in areas like RLHF, human-AI collaboration, interpretability, or control.
- Ph.D. or other advanced degree in computer science, machine learning, or a related field.
- 4+ years of research engineering experience.
- Proficiency in Python or similar programming languages.
- Enthusiasm for long-term AI safety and deep thought on technical paths to safe AGI.
- Willingness to apply methods from interpretability, robustness, alignment, and control to ensure model safety.
- Ability to thrive in environments involving large-scale AI systems.