Principal Research Engineer, Model Training & Post-Training
$400k - $550k • Palo Alto, California, United States
Posted 1mo ago
Job Location
Palo Alto, California, United States
Tech Stack
Remote Work Policy
On-site
Categories
AI Research Engineer
About the job
Inflection AI is seeking a hands-on technical leader to own the model-improvement loop, from data and training through evaluations, post-training, release criteria, and production feedback. This role sits at the intersection of research, production engineering, and model release, with the goal of shipping measurably better models for users. The ideal candidate will have prior experience leading significant model training or post-training initiatives and can make informed tradeoffs across data, compute, architecture, and quality to achieve a clear technical roadmap.
Responsibilities
- Own the model-improvement roadmap covering capability, reliability, emotional intelligence, tool use, safety, latency, cost, and enterprise readiness.
- Lead training and post-training strategy, including supervised fine-tuning, RLHF, DPO, GRPO, RLAIF, reward modeling, preference optimization, tool-use fine-tuning, distillation, synthetic data, and related methods.
- Drive model architecture and optimization decisions for modern transformer-based and hybrid architectures, focusing on both training-time and inference-time performance.
- Lead large-scale training efforts on distributed GPU clusters, including systems operating at the scale of 1,000+ GPUs.
- Define and execute data strategy encompassing data curation, mixture design, deduplication, decontamination, human-in-the-loop pipelines, preference data, evaluation data, synthetic data, and production feedback loops.
- Build and improve evaluation and release-quality systems, including model evaluations, quality gates, regression detection, release criteria, model-readiness reviews, and post-release monitoring.
- Partner with infrastructure and research engineering teams to enhance distributed training reliability, checkpointing, fault tolerance, observability, reproducibility, and cost-performance tradeoffs.
- Debug and improve model behavior across the full stack: data, training, post-training, evaluation, infrastructure, product integration, and production feedback.
Requirements
- Experience leading or serving as a principal contributor to large-scale LLM, multimodal, or foundation-model training or post-training programs.
- Deep experience with transformer-based models, hybrid architectures, modern deep-learning frameworks, and distributed training systems.
- Strong practical experience with post-training and alignment methods such as SFT, RLHF, DPO, GRPO, RLAIF, reward modeling, preference optimization, tool-use fine-tuning, or related approaches.
- Experience operating or partnering on large-scale training infrastructure, ideally including GPU clusters at the scale of 1,000+ GPUs.
- Strong systems instincts regarding throughput, cost, reliability, observability, debugging, checkpointing, reproducibility, and fault tolerance.
- Excellent judgment concerning data quality, evaluation design, model regressions, release readiness, and production model behavior.
- Ability to balance research ambition with product pragmatism, user impact, and operational discipline.
- Experience leading senior technical teams while continuing to contribute directly to technical decisions and implementation.
- PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related field, or equivalent practical experience.
Benefits
- Meaningful equity component
- Robust medical, dental and vision options with employer contributions for HSA, FSA and DFSA
- 401k matching program
- Flexible Time Off
- 10 paid holidays
- 5 days sick leave
- Parental, Medical and Family care leave
- Generous cell-phone, wellness and office set up stipends
- Support of country-specific visa needs for international employees living in the Bay Area