Member of Technical Staff, Data Flywheel
New York, NY • FullTime
Posted 19d ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Reflection is a research lab dedicated to making intelligence open and accessible. We build open models that empower users to control their intelligence and shape the future of AI. The Data Flywheel team plays a crucial role in bridging the gap between theoretical benchmark performance and practical, real-world utility. We achieve this by identifying and developing the signals, data, and feedback mechanisms that transform model usage into rigorous evaluations, targeted training data, and measurable advancements for future AI generations. This is a hands-on technical position situated at the nexus of research and deployment, where you will guide ambiguous model behaviors from initial observation through measurement, intervention, and validated improvement. You will collaborate across evaluation, human and synthetic data, infrastructure, post-training, and live deployment aspects, working closely with internal researchers and engineers, as well as external customers, partners, vendors, and the open-source community.
Responsibilities
- Identify high-value data sources and partnership opportunities, understanding underlying use cases to translate them into representative evaluations.
- Integrate new data sources, managing partner conversations, data scoping, quality validation, and integration into production evaluation and training pipelines.
- Design and build evaluations, graders, and feedback loops to quantify priority real-world model behaviors.
- Analyze model performance and failure modes to inform targeted datasets, reward signals, and training interventions.
- Develop human and synthetic data strategies for capabilities lacking sufficient existing data, including managing vendor evaluation and data collection programs.
- Build and maintain the infrastructure and pipelines for reliable, large-scale data ingestion, inspection, versioning, and evaluation.
- Collaborate with pre-training, post-training, applied, and partnership teams to translate new signals into measurable model improvements.
Requirements
- Bachelor's, Master's, or PhD in Computer Science, Machine Learning, or a related field, or equivalent practical experience.
- Deep technical understanding of LLM training and evaluation, with hands-on experience in areas like evaluation design, data curation, reinforcement learning, or reward design.
- Strong software engineering skills and experience building automated data/evaluation pipelines or large-scale ML systems.
- Proven track record of owning high-impact projects end-to-end, navigating ambiguity, and adapting to changing priorities.
- Highly collaborative and action-oriented approach, with enthusiasm for defining how a new frontier lab measures and accelerates model progress.
- High agency and ability to thrive in a fast-paced startup environment, with a bias for impact over process.
- Enjoyment of collaboration across research, engineering, operations, and product disciplines.
Benefits
- Top-tier compensation (salary and equity)
- Stock options
- Comprehensive medical, dental, vision, and life insurance
- Annual wellness allowance
- Provided lunch and dinner in the office
- 22 weeks paid parental leave
- Unlimited paid time off (U.S.)
- 30 days paid time off (U.K.)
- Visa sponsorship support
- Regular off-sites, happy hours, and team celebrations