Research Engineer, Safety & Alignment

Remote • San Francisco • FullTime

Posted 5h ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

AI Agent Engineer

About the job

We are an applied AI lab building end-to-end software agents, notably the AI software engineer, Devin. Our team comprises top talent from leading AI companies and competitive programming backgrounds. We are seeking a Research Engineer to join our founding safety team. This role is crucial for ensuring the safe and correct behavior of our AI agents in real-world applications. You will be instrumental in developing evaluations, red-teaming strategies, and alignment methods that directly influence how our models are trained and deployed, ensuring they operate reliably and ethically.

Responsibilities

  • Design and run evaluations for autonomous agent failure modes such as unsafe actions, prompt injection, data exfiltration, sandbox escape, reward hacking, and misuse, ensuring they are fast enough for every release.
  • Systematically attack the agent and its harness through red-teaming to identify failures before customers do, converting them into fixes and regression tests.
  • Develop alignment methods, including post-training techniques like reward modeling, preference data, and constitutional approaches, to enhance model safety without compromising performance.
  • Collaborate with the agent team on permissions, oversight, and escalation mechanisms to define when and how the agent should act, ask for help, or explain its actions.
  • Contribute to external safety research through publications, evaluations, and open methods where appropriate.

Requirements

  • Hands-on experience in safety, alignment, or interpretability research, with a track record of building systems that influenced model training or deployment.
  • Understanding of agentic systems and practical failure modes like tool misuse, specification gaming, long-horizon drift, and adversarial inputs.
  • Proficiency in Python and PyTorch (or JAX), with the ability to build eval infrastructure, run experiments at scale, and understand training/inference code.
  • Demonstrated empirical rigor in designing experiments, reporting results, and distinguishing real improvements from noise.
  • Comfort and skill in adversarial thinking and system breaking.
  • Relevant industry experience in a frontier AI lab, applied AI company, or developer tools company.
  • An advanced degree (PhD) in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline, or equivalent industry research experience.

About Cognition

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.