Researcher, Safety Training, National Security
San Francisco • FullTime
Posted 2d ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Research Engineer
About the job
We are seeking a researcher to train and evaluate models for U.S. government use, with a focus on national security applications. You will advance safety post-training and robustness, helping models follow nuanced policies while preserving their usefulness and capabilities. This role involves researching and implementing methods for safety training, reinforcement learning, and adversarial robustness, as well as developing evaluations, identifying model failure modes, and using findings to improve training. You will also collaborate with research, engineering, security, and policy partners to support safe, reliable deployment.
Responsibilities
- Research and implement methods for safety training, reinforcement learning, and adversarial robustness.
- Develop evaluations, identify model failure modes, and use findings to improve training.
- Work with research, engineering, security, and policy partners to support safe, reliable deployment.
Requirements
- 4+ years of relevant AI safety research experience, including RLHF, adversarial training, or robustness.
- Degree in computer science, machine learning, or a related field.
- Strong deep learning research or engineering skills.
- Experience improving model safety for deployment.
- Enjoy collaborative research.
- Motivated by OpenAI’s mission and the responsible use of AI in safety-critical settings.
- Active TS/SCI clearance or equivalent.