Model Policy Manager, Agentic Safety

San Francisco FullTime

Posted 2h ago

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

This role focuses on understanding and mitigating real-world risks arising from model misalignment as AI models become more autonomous and operate over longer timeframes. You will investigate how misaligned behavior emerges across extended interactions, such as models pursuing incorrect objectives, taking unsafe shortcuts, or circumventing constraints. The insights gained will be translated into behavioral policies, evaluations, monitoring systems, and safeguards to improve and validate model behavior. This position is ideal for individuals passionate about transforming alignment and safety concerns into concrete, empirically validated improvements for advanced AI systems.

Responsibilities

  • Identify vulnerabilities in model interactions with tools, data, and external systems, and develop corresponding safeguards.
  • Develop threat models and empirical frameworks to understand harmful outcomes from misaligned AI behavior.
  • Build frameworks for analyzing harmful outcomes stemming from model misalignment.
  • Identify the underlying behaviors and system conditions that lead to harmful outcomes.
  • Translate findings into policy frameworks, evaluation criteria, online measurement strategies, and safeguards.
  • Develop human data campaigns and gold sets to measure and evaluate emerging behaviors and risks.
  • Collaborate with research, engineering, security, and product teams to enhance model and system safety, balancing safety, utility, and business risk.
  • Inform deployment decisions, system cards, safeguard reports, and OpenAI's overall approach to agentic safety.
  • Develop monitoring approaches to detect regressions and emerging risks post-deployment.

Requirements

  • Strong background in AI agent safety, privacy, security, cybersecurity, or related fields, with an adversarial mindset.
  • Demonstrated interest in AI alignment and a solid understanding of the technical drivers of misaligned model behavior.
  • Technical fluency to work with evaluation and training data, analyze data patterns, and identify limitations.
  • Comfort working hands-on with model data and evaluation results, including inspecting examples and analyzing failure patterns.
  • Ability to use empirical evidence to develop and refine safety policies and safeguards.
  • Capacity to translate complex or ambiguous alignment risks into precise behavioral expectations and measurable evaluation criteria.
  • Effectiveness in working across research, engineering, security, product, and policy teams.
  • Clear communication skills for discussing complex and uncertain technical risks.
  • Adaptability to fast-paced, collaborative research environments with shifting priorities.

Benefits

  • Relocation support for new employees
  • Hybrid work model (three days in office, optional WFH Thursdays/Fridays)
  • Height-adjustable desks
  • Conference rooms
  • Phone booths
  • Well-stocked kitchens with snacks and drinks
  • Three in-house prepared meals daily
  • Private outdoor space
  • Nap rooms
  • Private bike storage

About OpenAI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.