Post-Training Applied Researcher
$200k - $275k • Remote • San Francisco • FullTime
Posted 22h ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Research Engineer
About the job
Baseten powers mission-critical inference for leading AI companies, enabling them to bring cutting-edge models into production. We are seeking a Post-Training Applied Researcher to work directly with stakeholders from fast-growing AI companies. In this role, you will focus on post-training open-source models to excel on specialized tasks. Your day-to-day will involve extracting signal from complex datasets and building the necessary components like reward functions, environments, and training pipelines to improve models. The models you train will be deployed to production and used by millions of users. We are looking for individuals with hands-on LLM fine-tuning and RL experience who are eager to ship models, translate customer requirements into training curricula, and balance rigor with rapid iteration.
Responsibilities
- Design and execute post-training pipelines including SFT, GRPO, DPO, RLVR, reward function engineering, and synthetic data generation.
- Develop task-specific training environments and evaluation harnesses tailored to customer domains such as healthcare, code generation, and legal, covering multi-turn tool use, sandboxed execution, and agentic workflows.
- Collaborate with customers to transform production data into training signals, design reward loops based on real usage patterns, and manage distribution shift.
- Conduct and analyze end-to-end training experiments, diagnosing issues like reward hacking, importance sampling drift, and advantage estimation instabilities.
- Publish research findings in top-tier venues and contribute to Baseten's open-source training libraries.
Requirements
- Hands-on experience training LLMs with reinforcement learning, with a demonstrated understanding of GRPO or PPO beyond basic implementation, including group advantage computation, clipped objectives, and KL penalty design.
- Strong intuition for reward engineering, capable of distinguishing effective rewards from those prone to exploitation.
- Experience building multi-turn agent environments with tool use, extending beyond single-turn question-answering.
- Proficiency across the entire ML pipeline, from dataset construction to training, evaluation, and deployment.
- Experience with production ML systems, with a preference for candidates who have implemented a training-inference loop where production data informs model improvement.
- Experience with RL training frameworks.
- Publications at top AI conferences (NeurIPS, ICML, ICLR) focused on RL for LLMs, reward modeling, or alignment.
Benefits
- Competitive compensation and meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employee and dependents (U.S. only).
- Flexible PTO policy with company-wide Winter Break.
- Paid parental leave.
- Fertility and family-building stipend through Carrot.
- Company-facilitated 401(k) (U.S. only).
- Exposure to a variety of ML startups for learning and networking opportunities.