Forward Deployed Engineer - LLM Post-training

San Francisco, CA FullTime

Posted 3mo ago

Job Location

San Francisco, CA

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

LLM Engineer

About the job

Reflection is a research lab dedicated to making intelligence open and accessible. We build open-weight models that empower users to control their AI and shape its future. As a core member of the Applied AI team, you will drive model fine-tuning and evaluations for enterprise customers. This role involves adapting our open-weight models for specific customer domains, tasks, and constraints, working hands-on with customer data, running fine-tuning workflows, building evaluation harnesses, and deploying adapted models to production. You will collaborate directly with customers to understand their needs and with research teams to advance the possibilities of AI.

Responsibilities

  • Fine-tune open-weight models for customer-specific use cases, including dataset preparation, configuration of training runs (SFT, preference optimization, reinforcement fine-tuning), and iteration based on evaluations.
  • Build and maintain evaluation infrastructure, including designing eval suites, curating test sets, establishing baselines, and measuring model performance on customer-relevant tasks.
  • Prepare training data from raw customer inputs, involving data quality inspection, cleaning, formatting, identification of adversarial or noisy samples, and building reproducible data pipelines.
  • Debug and diagnose training and inference issues, interpreting loss curves, identifying data quality problems, and recognizing training dynamics that indicate issues.
  • Support end-to-end deployments of fine-tuned models across hybrid environments (public cloud, VPC, on-premises), ensuring inference performance and reliability.
  • Contribute to evolving playbooks, evaluation benchmarks, and best practices within the fine-tuning and evaluations practice.

Requirements

  • Applied ML experience with hands-on fine-tuning of language models, including dataset preparation, running training loops, evaluating results, and shipping fine-tuned models.
  • Familiarity with SFT, DPO, RLHF, or similar fine-tuning techniques.
  • Understanding of evaluation methodology, including designing evaluations, interpreting training graphs, and assessing model improvement beyond benchmark overfitting.
  • Comfort with training infrastructure such as GPUs and compute management, and debugging common training failures.
  • Strong software engineering fundamentals in Python, with experience writing clean, reproducible code.
  • Experience with data pipelines and version control for datasets and experiments.
  • 3+ years of engineering experience with significant exposure to applied ML or ML engineering.
  • Demonstrated ability and interest in customer-facing environments, translating user needs and domain requirements into training strategies.
  • Self-starter with high agency and ownership, capable of thriving in fast-paced startup environments.

Benefits

  • Top-tier compensation (salary and equity)
  • Stock options
  • Comprehensive medical, dental, vision, and life insurance
  • Annual wellness allowance
  • Provided lunch and dinner in the office
  • 22 weeks paid parental leave
  • Unlimited paid time off (U.S.)
  • 30 days paid time off (U.K.)
  • Visa sponsorship support
  • Regular off-sites, happy hours, and team celebrations

About Reflection ai

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.