Model Policy Manager
Remote • San Francisco • FullTime
Posted 1d ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Frontier AI systems present both immense opportunities and significant safety challenges. This role focuses on defining how OpenAI's models should behave in high-risk or ambiguous situations, including agentic and multimodal systems, user safety, and privacy. The ideal candidate will be adept at navigating unfamiliar topics, reasoning from first principles, and translating ambiguity into practical, enforceable model behaviors. You will collaborate closely with research, engineering, product, preparedness, and operations teams to develop technically grounded, measurable policies that are responsive to real-world risks.
Responsibilities
- Design and maintain model policies for safety-relevant domains, including dual-use, agentic, and emerging frontier-risk areas.
- Translate risk and harm models into clear behavioral specifications, evaluation criteria, grading guidance, and system-level safeguards.
- Define practical boundaries between beneficial AI uses and assistance that could enable harm, exploitation, misuse, or unsafe outcomes.
- Build policy artifacts to support model training, evaluation, and deployment.
- Partner with safety researchers, engineers, product teams, and stakeholders to operationalize policy into scalable model behavior and measurable safeguards.
- Utilize red-teaming results, deployment data, model failures, and edge cases to improve policy and evaluation quality.
- Identify emerging capabilities of frontier AI systems that could create new safety challenges or lower barriers to harm.
- Study real-world deployments to identify where model behavior succeeds, fails, or drifts from intended safety posture.
- Combine long-horizon safety research with hands-on launch and deployment work.
- Contribute to system cards, safety reports, policy documentation, launch reviews, and external communications on model safety and risk mitigation.
- Design and run human data campaigns, including gold set construction, labeling guidance, and adjudication, to ensure policies can be reliably measured and improved.
Requirements
- Strong judgment regarding the impact of advanced AI systems on real-world risk, especially in ambiguous or high-impact areas.
- Experience building or applying policies, taxonomies, harm models, threat models, or risk frameworks for complex technical, social, or adversarial systems.
- Ability to work across domains without being the deepest subject-matter expert, while knowing when to seek expert input.
- Capacity to transform ambiguous questions into structured policy frameworks, evaluation criteria, operational guidance, and enforceable model behavior.
- Comfort using empirical evidence, including evaluations, red-teaming results, and deployment observations, to inform policy decisions.
- Systems thinking across policy, data, graders, classifiers, training, deployment safeguards, measurement, monitoring, and escalation workflows.
- Technical judgment on what model behavior can realistically be trained, measured, evaluated, and enforced at scale.
- Ability to work effectively across research, engineering, product, policy, domain experts, and operational teams.
- Clear writing skills for complex tradeoffs involving safety, user value, and implementation constraints.
- Pragmatic approach to safety, focused on reducing real-world risk while preserving legitimate, beneficial, and socially valuable AI uses.
- Enjoyment of fast-paced, collaborative research environments with shifting priorities.
- Grounded in implementation details, empirical results, and what can actually be trained or measured.
Benefits
- Relocation support to new employees