Research Engineer, Safety & Alignment
Remote • San Francisco • FullTime
Posted 5h ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Agent Engineer
About the job
We are an applied AI lab building end-to-end software agents, notably the AI software engineer, Devin. Our team comprises top talent from leading AI companies and competitive programming backgrounds. We are seeking a Research Engineer to join our founding safety team. This role is crucial for ensuring the safe and correct behavior of our AI agents in real-world applications. You will be instrumental in developing evaluations, red-teaming strategies, and alignment methods that directly influence how our models are trained and deployed, ensuring they operate reliably and ethically.
Responsibilities
- Design and run evaluations for autonomous agent failure modes such as unsafe actions, prompt injection, data exfiltration, sandbox escape, reward hacking, and misuse, ensuring they are fast enough for every release.
- Systematically attack the agent and its harness through red-teaming to identify failures before customers do, converting them into fixes and regression tests.
- Develop alignment methods, including post-training techniques like reward modeling, preference data, and constitutional approaches, to enhance model safety without compromising performance.
- Collaborate with the agent team on permissions, oversight, and escalation mechanisms to define when and how the agent should act, ask for help, or explain its actions.
- Contribute to external safety research through publications, evaluations, and open methods where appropriate.
Requirements
- Hands-on experience in safety, alignment, or interpretability research, with a track record of building systems that influenced model training or deployment.
- Understanding of agentic systems and practical failure modes like tool misuse, specification gaming, long-horizon drift, and adversarial inputs.
- Proficiency in Python and PyTorch (or JAX), with the ability to build eval infrastructure, run experiments at scale, and understand training/inference code.
- Demonstrated empirical rigor in designing experiments, reporting results, and distinguishing real improvements from noise.
- Comfort and skill in adversarial thinking and system breaking.
- Relevant industry experience in a frontier AI lab, applied AI company, or developer tools company.
- An advanced degree (PhD) in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline, or equivalent industry research experience.