Research Scientist - Multimodal Agent, Consumer Devices

Remote San Francisco FullTime

Posted 1mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

AI Agent Engineer

About the job

We are seeking a Research Engineer / Scientist to join our applied research team focused on developing new methods, models, and evaluation frameworks for the future of computing. This role will concentrate on building the learning and evaluation foundations that enable AI models to become more context-aware, adaptive, and useful over time. You will tackle challenges such as reward modeling, preference learning, long-horizon evaluation, and policy improvement for systems that require high-quality behavioral decisions in realistic user settings. The work is deeply product-grounded, aiming for improved model behavior in real-world use rather than just benchmark performance. The ideal candidate is passionate about advancing AI systems beyond simple assistant interactions towards those that improve through feedback, learn from richer signals, and are trained against meaningful notions of user value.

Responsibilities

  • Develop RLHF and post-training methods for multimodal models.
  • Build reward models and preference-learning pipelines for adaptive, personalized model behavior.
  • Design datasets, rubrics, and evaluation frameworks to capture user preferences, contextual appropriateness, and long-term value in realistic tasks.
  • Run experiments on policy improvement using explicit feedback, implicit signals, and model-based grading.
  • Address long-horizon evaluation problems where model quality depends on cumulative behavior improvement over time.
  • Collaborate with safety researchers to ensure adaptation and personalization remain aligned, interpretable, and bounded by clear constraints.
  • Prototype and iterate quickly on training recipes, reward formulations, data pipelines, and evaluation suites for product-relevant behaviors.
  • Help define how success is measured for personalized AI systems, including trust, appropriateness, and long-term user benefit.

Requirements

  • Strong background in machine learning research, with experience in RLHF, reward modeling, preference optimization, or post-training for large models.
  • Experience in reinforcement learning, ranking, recommender systems, personalization, memory, or human-in-the-loop evaluation.
  • Proficiency in designing clean experiments, reliable evaluations, and decision-useful metrics.
  • Excitement for training models against nuanced behavioral objectives.
  • Experience building datasets or eval pipelines grounded in human preferences, rubrics, or real-world product behavior.
  • Comfort working across the full stack, from data generation and labeling strategy to training runs, reward functions, and analysis.
  • Interest in multimodal AI and how models learn from richer interaction signals over time.
  • Desire to work on product-shaping research with high stakes for trust, alignment, and long-term user value.
  • Enjoyment of close collaboration with engineers, designers, and safety researchers.

Benefits

  • Relocation assistance to new employees.

About OpenAI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.