Research Engineer, Safety
$200k - $400k • San Francisco • FullTime
Posted 9d ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Research Engineer
About the job
Decagon is seeking a Research Engineer focused on Safety to ensure the reliability and controllability of their AI agents from evaluation through production. This role involves identifying real-world failure modes and developing the necessary models, evaluations, and safeguards to prevent them. The ideal candidate is a strong engineer passionate about advancing applied AI safety in production, with the autonomy to own their work end-to-end, ship impactful improvements, and make high-stakes technical decisions.
Responsibilities
- Research and build safeguards against prompt injection, unsafe tool use, sensitive-data disclosure, policy violations, and hallucinated commitments.
- Build adversarial evaluations, simulations, red-team datasets, and regression suites informed by production failures.
- Develop and deploy classifiers, judges, reward signals, post-training methods, and runtime safeguards for safer agent behavior.
- Analyze production traces and incidents to identify root causes, test mitigations, and measure their impact.
- Partner with Security, Product, Infrastructure, Legal, and customer-facing teams to translate enterprise requirements into scalable safeguards and rollout practices.
Requirements
- 2+ years of experience in AI/ML engineering, research, or AI safety.
- Hands-on experience evaluating, post-training, or deploying language models or agentic systems.
- Experience with modern post-training techniques such as reinforcement learning, preference optimization, distillation, model routing, and synthetic-data generation.
- Experience with adversarial testing, model red teaming, prompt injection, policy enforcement, privacy, or safe tool use.
- Fluency in Python and modern ML tooling, with strong experimental judgment and engineering depth to ship production systems.
- Comfort owning ambiguous, high-stakes technical problems and making clear risk and product tradeoffs.
- Experience building safeguards for high-stakes or regulated enterprise workflows (preferred).
- Familiarity with human-in-the-loop review, incident response, or responsible rollout frameworks for ML systems (preferred).
Benefits
- Equity
- Medical, Dental, and Vision benefits for you and your family
- Life Insurance and Disability Benefits
- Retirement Plan (e.g., 401K, pension)
- Parental Leave
- Fertility and family building benefits through Carrot
- Monthly stipend to support wellness, lifestyle, and work-life balance
- Daily lunches and snacks in the office