Agent Post-Training, Personality
San Francisco • FullTime
Posted 2mo ago
About the job
The Agent Post-training Personality team at OpenAI is responsible for shaping the collaborative capabilities of frontier AI agents. This role focuses on understanding and enhancing agent "personality," which encompasses thoughtfulness, clarity, perceptiveness, appropriate proactivity, and ease of interaction. The goal is to create agents that are not only effective but also exceptional collaborators, understanding user intent, communicating with good judgment, adapting to context, and taking initiative when needed. This position involves a blend of behavioral research, product thinking, and communication expertise, requiring close collaboration with product teams, human experts, and researchers across the organization to ensure these improvements are integrated into widely used AI models.
Responsibilities
- Develop a deep understanding of what constitutes a great collaborator for agents across various professional and everyday contexts.
- Translate qualitative judgments of model behavior into concrete hypotheses, evaluation metrics, graders, and training interventions.
- Analyze explicit and implicit user signals to identify behaviors that foster trust, satisfaction, continued use, and successful outcomes.
- Collaborate with human experts and trainers to generate high-quality preference data and rollouts that exemplify excellent collaborative behavior.
- Enhance reward models and reinforcement learning objectives to guide desired model behaviors.
- Work with pretraining and early-training teams on data mixtures, objectives, and synthetic data to influence downstream personality.
- Build and maintain sustainable pipelines for updating training data as understanding of ideal model behavior evolves.
- Partner with product teams like ChatGPT and Codex to translate consumer insights into model improvements and validate them in real-world workflows.
- Manage projects from initial observation of behavioral issues through experimentation, training, evaluation, and final launch.
Requirements
- Instinctively think from the user's perspective and prioritize how models feel to interact with, beyond benchmark performance.
- Translate subjective product questions into falsifiable hypotheses and rigorous evaluations while preserving nuance.
- Value individuality, adaptability, and behavioral diversity over optimizing for a single narrow style.
- Possess strong technical foundations in machine learning, software engineering, statistics, behavioral science, HCI, or a related field, with the ability to learn quickly across different parts of the technology stack.
- Demonstrate strong taste for model behavior, able to articulate why certain responses are thoughtful, natural, and useful.
- Experience with LLMs, post-training, RL/RLHF, reward modeling, evals, synthetic data, pretraining data, or production ML systems.
- Comfort working on ambiguous capability problems with noisy signals and qualitative failures, potentially requiring a combination of data, training, evals, and product changes.
- Ability to work effectively with diverse teams including researchers, engineers, product managers, designers, domain experts, and human-data teams, communicating clearly with each group.
- Willingness to build essential systems and processes, even if not glamorous.
- Desire to train and ship models that make agents genuinely useful for developers, enterprises, researchers, and everyday users.