AI Researcher, Core ML (Turbo)

$200k - $280k Remote San Francisco

Posted 1mo ago

Remote Work Policy

Fully remote

Categories

AI Research Engineer

About the job

The Turbo team operates at the intersection of efficient inference (algorithms, architectures, engines) and post-training/RL systems. We are responsible for building and managing the systems that power Together's API, focusing on high-performance inference and RL/post-training engines capable of operating at production scale. Our core mission is to advance the frontiers of efficient inference and RL-driven training, aiming to make models significantly faster and more cost-effective to run, while simultaneously enhancing their capabilities through RL-based post-training methods. This role involves working across the entire stack, from RL algorithms and training engines to kernels and serving systems, to develop and refine state-of-the-art models using RL pipelines. We value individuals with deep expertise in one area and a strong willingness to collaborate and grow across others.

Responsibilities

  • Advance inference efficiency end-to-end by designing and prototyping algorithms, architectures, and scheduling strategies for low-latency, high-throughput inference.
  • Implement and maintain changes in high-performance inference engines, including kernel backends, speculative decoding, and quantization.
  • Profile and optimize performance across GPU, networking, and memory layers to improve latency, throughput, and cost.
  • Unify inference with RL/post-training by designing and operating RL and post-training pipelines.
  • Make RL and post-training workloads more efficient with inference-aware training loops.
  • Use these pipelines to train, evaluate, and iterate on frontier models.
  • Co-design algorithms and infrastructure for tightly coupled objectives, rollout collection, and efficient inference.
  • Run ablations and scale-up experiments to understand trade-offs between model quality, latency, throughput, and cost.
  • Own critical systems at production scale by profiling, debugging, and optimizing inference and post-training services.
  • Drive roadmap items requiring real engine modification, including kernels, memory layouts, scheduling logic, and APIs.
  • Establish metrics, benchmarks, and experimentation frameworks to rigorously validate improvements.

Requirements

  • Deep expertise in at least one of the following: large-scale inference systems (e.g., SGLang, vLLM, FasterTransformer, TensorRT), GPU performance, distributed serving, RL/post-training for LLMs (e.g., GRPO, RLHF/RLAIF, DPO), reward modeling, Transformer architecture design, or distributed systems/HPC for ML.
  • Comfort working from algorithms to engines, with strong coding ability in Python.
  • Experience profiling and optimizing performance across GPU, networking, and memory layers.
  • Ability to translate new sampling methods, schedulers, or RL updates into production-grade implementations.
  • Solid research foundation with a track record of impactful work in ML systems, RL, or large-scale model training (papers, open-source projects, or production systems).
  • Ability to read new RL/post-training papers, understand their implications, and design minimal, correct changes.
  • Full-stack problem-solving skills, including identifying bottlenecks across the stack.
  • Enjoy collaborating with infra, research, and product teams, with a focus on both scientific quality and user-visible wins.
  • 3+ years of experience working on ML systems, large-scale model training, inference, or adjacent areas (or equivalent research/open-source experience).
  • Advanced degree in Computer Science, EE, or a related field, or equivalent practical experience.
  • Demonstrated experience owning complex technical projects end-to-end.

About Together AI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.