Research Program Manager - Model Evals and Safety
New York, NY • FullTime
Posted 3mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Research Engineer
About the job
Reflection is a research lab dedicated to making intelligence open and accessible. This role is foundational, focusing on building model evaluation and safety functions from the ground up. The Research Program Manager will be a key leader, driving clarity, decision-making, and coherence across teams. This is a 0-to-1 opportunity to define evaluation frameworks, build operational infrastructure for model safety, and establish processes that integrate safety into the model development lifecycle.
Responsibilities
- Build foundational infrastructure for model evaluations and safety, defining frameworks, tooling, and operational processes.
- Establish model safety operations, including workflows, review cadences, and decision frameworks.
- Partner with research and engineering leads to embed safety and evaluation checkpoints into the development process.
- Drive the scoping and prioritization of evaluation science and infrastructure investments.
- Establish engagement with the external safety ecosystem, including third-party assessments and academic partnerships.
- Create visibility and reporting structures for leadership on model safety status and risks.
- Champion a culture of blameless post-mortems and continuous learning for safety improvements.
Requirements
- 7+ years of experience in technical program management, research operations, or ML engineering, with experience building new functions from scratch.
- Familiarity with model evaluation and AI safety landscapes, including methodologies, red-teaming, and alignment research.
- Deep technical understanding to engage with researchers and engineers on model behavior, evaluation design, and system architecture.
- Proven ability to build functioning programs from ambiguous mandates.
- Strong stakeholder management skills with technical ICs, research leadership, and external partners.
- Excited to build from zero to one in a fast-moving team.
- Motivated by enabling responsible development of AI systems.
Benefits
- Top-tier compensation
- Stock options
- Comprehensive medical, dental, vision, and life insurance
- Annual wellness allowance
- Provided lunch and dinner in the office
- 22 weeks paid parental leave
- Unlimited paid time off (US) / 30 days paid time off (UK)
- Visa sponsorship support
- Regular off-sites, happy hours, and team celebrations