Forward Deployed Engineer, Lead - LLM Post-training
New York, NY • FullTime
Posted 9mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
LLM Engineer
About the job
Reflection is a research lab dedicated to making intelligence open and accessible. We are seeking an exceptional technical leader to build and scale our post-training and evaluation capabilities within the Applied AI team. This role involves taking our open-weight models and adapting them for specific customer domains, tasks, and constraints. You will own the end-to-end technical strategy for model customization, from synthetic data generation and reward modeling through training and production deployment, working directly with customers and research teams.
Responsibilities
- Lead post-training engagements with enterprise customers, including data assessment, training strategy definition, reward signal design, dataset preparation, training execution, and evaluation against benchmarks.
- Design and build RL training environments for model adaptation, including synthetic data generation pipelines, reward model training, and preference data collection workflows.
- Design and build evaluation infrastructure to define 'better' for customer use cases, build eval harnesses, curate test sets, and establish baselines for real-world performance.
- Own the data pipeline from raw customer data to training-ready datasets, encompassing synthetic data generation, quality inspection, cleaning, and standardization.
- Deploy post-trained models across hybrid environments (public cloud, VPC, on-premises), ensuring inference performance, cost efficiency, and scalability.
- Shape and scale the post-training and evaluation practice by defining playbooks, best practices, and technical standards, while mentoring engineers.
Requirements
- Hands-on post-training experience with large language models at scale, including building and operating RL training environments, designing preference optimization workflows on models 50B+ parameters, and shipping results to production.
- Experience building synthetic data generation pipelines, reward models, and verifiers for reinforcement learning workflows, including architecting data and feedback loops.
- Deep understanding of evaluation methodology for measuring performance, interpreting training dynamics, and distinguishing benchmark performance from real-world effectiveness.
- Practical experience with training infrastructure at scale, including multi-node GPU clusters, managing large training runs, debugging distributed training, and cost optimization.
- Strong software engineering fundamentals with experience in production-quality code, data pipelines, version control for datasets and models, and reproducible workflows.
- 6+ years of engineering experience, with at least 2 years focused on LLM post-training in a leadership capacity.
- Experience in customer-facing technical roles or a strong interest in developing this skill, with the ability to translate domain requirements into training strategies.
- Self-starter with high agency and ownership, capable of thriving in fast-paced startup environments.
Benefits
- Top-tier compensation (salary and equity)
- Stock options
- Comprehensive medical, dental, vision, and life insurance
- Annual wellness allowance
- Provided lunch and dinner in the office
- 22 weeks paid parental leave
- Unlimited paid time off in the U.S.
- 30 days paid time off in the U.K.
- Sponsorship support for visas and long-term immigration pathways
- Regular off-sites, happy hours, and team celebrations