Forward Deployed Engineer (Inference & Post-Training)
$270k - $300k • Remote • San Francisco
Posted 1mo ago
Job Location
San Francisco
Tech Stack
Remote Work Policy
Fully remote
Categories
LLM Engineer
About the job
As a Forward Deployed Engineer (FDE) focused on Inference & Post-Training, you will be a hands-on technical partner to strategic customers, assisting production AI teams with leveraging high-quality models and performing inference at scale. You will act as a deep-domain specialist in inference optimization, fine-tuning pipelines, and production deployment, partnering with Solutions Architects. FDEs add significant value by ensuring complex Proofs of Concept (POCs) are met, facilitating platform adoption, and guiding tailored optimization efforts, directly impacting customer success and company growth.
Responsibilities
- Optimize inference engines based on hardware, model architecture, and workload profiles.
- Develop configuration updates and tune performance for critical POCs, benchmarks, and customer deployments, including KV cache, speculative decoding, tensor parallelism, and quantization.
- Lead hands-on RL training runs and optimize system design for post-training and fine-tuning pipelines (LoRA, SFT, DPO, RLHF, GRPO).
- Serve as the primary technical contact for strategic accounts, monitoring and optimizing endpoint configurations and ensuring customers maximize platform value.
- Establish direct alignment with strategic customers during onboarding to ensure optimal inference and post-training configurations from the start.
- Provide product feedback by surfacing field insights to influence the software and model roadmap, and drive early feature adoption.
Requirements
- 5+ years of experience in a technical role with a focus on inference systems, open-source LLM deployment, or post-training workflows.
- Expert-level, hands-on experience with inference engines (e.g., vLLM, TensorRT-LLM, SGLang) and ability to diagnose performance issues.
- Deep knowledge of inference optimization techniques including KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization.
- Hands-on experience with fine-tuning and post-training pipelines (LoRA, SFT, DPO, RLHF, GRPO) and ability to advise on system design.
- Broad knowledge of state-of-the-art open-source models and strong judgment for model selection.
- Strong Python skills and comfort working in production environments.
Benefits
- Competitive compensation
- Startup equity
- Health insurance
- Flexibility in terms of remote work