Forward Deployed Engineer (Inference & Post-Training)

$270k - $300k Remote San Francisco

Posted 1mo ago

Remote Work Policy

Fully remote

Categories

LLM Engineer

About the job

As a Forward Deployed Engineer (FDE) focused on Inference & Post-Training, you will be a hands-on technical partner to strategic customers, assisting production AI teams with leveraging high-quality models and performing inference at scale. You will act as a deep-domain specialist in inference optimization, fine-tuning pipelines, and production deployment, partnering with Solutions Architects. FDEs add significant value by ensuring complex Proofs of Concept (POCs) are met, facilitating platform adoption, and guiding tailored optimization efforts, directly impacting customer success and company growth.

Responsibilities

  • Optimize inference engines based on hardware, model architecture, and workload profiles.
  • Develop configuration updates and tune performance for critical POCs, benchmarks, and customer deployments, including KV cache, speculative decoding, tensor parallelism, and quantization.
  • Lead hands-on RL training runs and optimize system design for post-training and fine-tuning pipelines (LoRA, SFT, DPO, RLHF, GRPO).
  • Serve as the primary technical contact for strategic accounts, monitoring and optimizing endpoint configurations and ensuring customers maximize platform value.
  • Establish direct alignment with strategic customers during onboarding to ensure optimal inference and post-training configurations from the start.
  • Provide product feedback by surfacing field insights to influence the software and model roadmap, and drive early feature adoption.

Requirements

  • 5+ years of experience in a technical role with a focus on inference systems, open-source LLM deployment, or post-training workflows.
  • Expert-level, hands-on experience with inference engines (e.g., vLLM, TensorRT-LLM, SGLang) and ability to diagnose performance issues.
  • Deep knowledge of inference optimization techniques including KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization.
  • Hands-on experience with fine-tuning and post-training pipelines (LoRA, SFT, DPO, RLHF, GRPO) and ability to advise on system design.
  • Broad knowledge of state-of-the-art open-source models and strong judgment for model selection.
  • Strong Python skills and comfort working in production environments.

Benefits

  • Competitive compensation
  • Startup equity
  • Health insurance
  • Flexibility in terms of remote work

About Together AI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.