Member of Technical Staff - Post Training
€130k - €340k • Freiburg (Germany) • FullTime
Posted 4mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
We are seeking a Member of Technical Staff specializing in Post Training to own the end-to-end post-training pipeline for our multimodal generative models. This role is crucial for transforming foundation models into polished products, encompassing data strategy, reward modeling, preference optimization, distillation, and safety tuning across image, editing, and video modalities. You will be instrumental in driving significant improvements in model quality, developing the infrastructure that accelerates research team iteration, and advancing the state-of-the-art in aligning generative models with human intent. This is a Staff/Senior individual contributor position for someone with proven experience shipping post-training for a frontier model.
Responsibilities
- Own the full post-training pipeline end to end, including data curation, reward modeling, fine-tuning, preference optimization, distillation, safety tuning, evaluation, and deployment.
- Advance techniques across the post-training stack such as SFT, RLHF, RLAIF, DPO, preference learning, and reward modeling to align models with human intent and aesthetic judgment.
- Work across various modalities including text-to-image, image editing, multi-reference, and video post-training.
- Build personalization and customization capabilities to allow users to adapt models to their creative styles.
- Design and maintain high-throughput fine-tuning and evaluation infrastructure to support rapid iteration.
- Identify and address quality and alignment gaps through rigorous evaluation and targeted research and engineering.
Requirements
- Proven experience owning post-training for a frontier generative model through release, including SFT, preference optimization (DPO or RLHF), distillation, and safety tuning, with measurable quality wins.
- Deep experience across the post-training stack, including reward modeling, preference learning, RLHF/RLAIF, and personalization.
- Comfort working across modalities: text-to-image, image editing, multi-reference, and ideally video.
- Strong PyTorch fluency and ability to write research code that is maintainable and buildable upon.
- Experience with distillation techniques (e.g., LADD, DMD, consistency models) or building high-throughput evaluation pipelines is a strong plus.
- A bias toward shipping measurable model-quality improvements that reach users.
Benefits
- Equity
- Compensation for reasonable travel costs