Researcher, Recursive Self-Improvement Safety
San Francisco • FullTime
Posted 1mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Research Engineer
About the job
OpenAI is seeking strong technical executors to join the Preparedness team, focusing on mitigating frontier risks associated with accelerated AI development, including recursive self-improvement. This role involves anticipating future misalignment risks and developing strategic solutions. Responsibilities span designing and implementing pre-deployment risk assessments, control measures, training interventions, and translating technical work into institutional practices and external communications. The work is urgent, fast-paced, and has significant implications for the future of AI development and society.
Responsibilities
- Anticipate and address future misalignment risks associated with AI development, particularly recursive self-improvement.
- Design and implement pre-deployment risk assessments and control measures for AI systems.
- Develop and implement RSI-relevant training interventions.
- Translate technical insights into established institutional practices and external-facing communications.
- Develop scalable oversight practices for monitoring model misbehavior in superhuman capability regimes.
- Create automated auditing approaches to identify severe forms of model misalignments.
- Conduct rigorous testing and red-teaming of model misbehavior measurements, focusing on loss-of-control risks.
- Design experiments and evaluations to understand model misalignment and safety-relevant capabilities.
- Prototype technical mechanisms for verifying compliance with future AI safety agreements.
- Track progress toward automation of technical staff to inform AI alignment and security investments.
- Identify and address blindspots in mitigation areas for RSI safety.
- Perform hypothesis-driven research and translate insights into interventions or control systems impacting production models.
- Collaborate with or manage other staff to rapidly scale efforts.
- Turn open-ended objectives into concrete, prioritized directions.
- Build and iteratively improve scrappy prototypes.
- Secure buy-in from other staff and communicate work clearly.
Requirements
- Exceptional technical executor.
- Strong strategic and research taste; ability to prioritize effectively in domains with weak feedback loops.
- Passion for mitigating risks associated with recursive self-improvement.
- Driven by a desire to positively impact the future of AI development.
- Experience in ML research, AI alignment, or AI verification is a bonus.