Member of Technical Staff - Safety
San Francisco, CA • FullTime
Posted 6mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Reflection is a research lab dedicated to making intelligence open and accessible. We build open models that empower users to control their intelligence and shape the future of AI. As a Member of Technical Staff - Safety, you will be instrumental in ensuring the safety and reliability of our AI models. This role involves owning the red-teaming and adversarial evaluation pipeline, translating safety findings into concrete guardrails, and validating that every release meets our risk thresholds before deployment. You will develop scalable, automated safety benchmarks and research state-of-the-art jailbreaking techniques and defenses to proactively address potential vulnerabilities.
Responsibilities
- Own the red-teaming and adversarial evaluation pipeline for Reflection’s models.
- Translate safety findings into concrete guardrails in collaboration with the Alignment team.
- Validate that every release meets the lab’s risk thresholds before shipping.
- Develop scalable, automated safety benchmarks that evolve with model capabilities.
- Research and implement state-of-the-art jailbreaking techniques and defenses.
Requirements
- Graduate degree (MS or PhD) in Computer Science, Machine Learning, or related discipline, or equivalent practical experience in AI Safety.
- Deep technical understanding of LLM safety, including adversarial attacks, red-teaming methodologies, and interpretability.
- Strong software engineering capabilities with experience building automated evaluation pipelines or large-scale ML systems.
- Experience with Reinforcement Learning (RLHF/RLAIF) and its impact on model safety and alignment is a strong plus.
- Ability to thrive in a fast-paced, high-agency startup environment with a bias toward action.
- Willingness to make high-stakes decisions regarding model release and safety thresholds.
- Passion for advancing the frontier of intelligence.
Benefits
- Top-tier compensation: Salary and equity structured to recognize and retain talent globally.
- Stock options.
- Comprehensive medical, dental, vision, and life insurance.
- Annual wellness allowance.
- Lunch and dinner provided in the office daily.
- 22 weeks paid parental leave for all new parents.
- Unlimited paid time off in the U.S.
- 30 days paid time off in the U.K.
- Visa sponsorship support.
- Regular off-sites, happy hours, and team celebrations.