Member of Technical Staff, Evals & Post-Training Product

Remote San Mateo FullTime

Posted 10mo ago

Job Location

San Mateo

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Fireworks is seeking a Member of Technical Staff, Evals & Post-Training Product to define how developers improve models on the Fireworks platform. This role combines scalable system design, deep data science, and model quality. You will build the infrastructure and workflows connecting evaluation and post-training, taking our evaluation setup to the next stage by improving programmatic access and scale. You will work across backend systems, sandbox infrastructure, and user-facing surfaces to simplify the authoring of evaluations, understanding of results, and rapid iteration.

Responsibilities

  • Scale Eval Infrastructure: Own and evolve the internal eval setup, designing systems to eliminate single-host coordination bottlenecks, resolve log-syncing latency, and build seamless programmatic access.
  • Benchmark Obsession & Reproduction: Track state-of-the-art benchmarks, read research papers, analyze data discrepancies, and rigorously reproduce published results.
  • Pioneer New Benchmarks: Design and build new benchmarks to measure model performance on complex, emerging, or domain-specific use cases.
  • Own Fine-Tuning Product Experiences: Build and improve user-facing workflows for post-training, including fine-tuning experiences across SFT, RFT, and related model-improvement capabilities.
  • Work Closely With Users: Partner with customers and internal stakeholders to understand evaluation and fine-tuning needs, triage issues, and convert bespoke workflows into productized, reusable solutions.

Requirements

  • 1–7 years of software engineering or data science experience.
  • Strong system design skills, including architecting scalable, programmatic systems and transitioning legacy setups.
  • Hands-on experience building or working with sandbox environments for secure code execution and testing.
  • Analytical and data science mindset with a deep understanding of LLM evaluations, their design, and using results to guide model improvement.
  • Meticulous about data and metrics.
  • Understanding of the end-to-end GenAI lifecycle, from prompting to productionizing agents.
  • 3+ years of software engineering or applied data science experience (preferred).
  • Experience with Harbor framework or similar container registry and orchestration tools (preferred).
  • Interest in discovering and publicly sharing insights on model performance (preferred).
  • Interest in AI hardware, GPU constraints, and inference optimization (preferred).
  • Startup DNA: Experience in fast-paced environments with end-to-end feature ownership (preferred).

Benefits

  • Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure.
  • Build What’s Next: Work with bleeding-edge technology impacting businesses and developers globally.
  • Ownership & Impact: Join a fast-growing team where your work directly shapes the future of AI.
  • Learn from the Best: Collaborate with world-class engineers and AI researchers.

About fireworks ai

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.