Research Engineer, Core ML

$200k - $280k Remote San Francisco

Posted 1mo ago

Remote Work Policy

Fully remote

Categories

AI Research Engineer

About the job

This research engineering role focuses on translating new Reinforcement Learning (RL) algorithms, scheduling methods, and inference optimizations into production-grade systems that power Together's API. The Core ML team operates at the intersection of efficient inference (algorithms, architectures, engines) and post-training/RL systems, building and maintaining high-performance inference and RL engines at production scale. The goal is to significantly improve model speed, cost-efficiency, and capabilities through RL-based post-training. This position requires a blend of algorithmic understanding and systems engineering, with opportunities to work across the entire stack from RL algorithms and training engines to kernels and serving systems, ultimately driving measurable improvements in latency, throughput, cost, and model quality at scale.

Responsibilities

  • Advance inference efficiency end-to-end by designing and prototyping algorithms, architectures, and scheduling strategies for low-latency, high-throughput inference.
  • Implement and maintain changes in high-performance inference engines, including kernel backends, speculative decoding, and quantization.
  • Profile and optimize performance across GPU, networking, and memory layers to improve latency, throughput, and cost.
  • Unify inference with RL/post-training by designing and operating RL and post-training pipelines.
  • Make RL and post-training workloads more efficient with inference-aware training loops.
  • Use RL pipelines to train, evaluate, and iterate on frontier models.
  • Co-design algorithms and infrastructure for tight coupling between objectives, rollout collection, evaluation, and efficient inference.
  • Run experiments to understand trade-offs between model quality, latency, throughput, and cost, and feed insights back into design.
  • Own critical systems at production scale by profiling, debugging, and optimizing inference and post-training services.
  • Drive roadmap items requiring engine modification, including kernels, memory layouts, scheduling logic, and APIs.
  • Establish metrics, benchmarks, and experimentation frameworks for rigorous validation.
  • Provide technical leadership by setting direction for cross-team efforts and mentoring other engineers and researchers.

Requirements

  • Bias toward implementation and shipping, with excitement to modify real engines and services.
  • Strong expertise in at least one of the following: large-scale inference systems, GPU performance, distributed serving, RL/post-training for LLMs, reward modeling, Transformer architecture design, or distributed systems/HPC for ML.
  • Comfort working from algorithms to engines, including strong coding ability in Python.
  • Experience profiling and optimizing performance across GPU, networking, and memory layers.
  • Ability to translate new sampling methods, schedulers, or RL updates into production-grade implementations.
  • Solid research foundation in areas of depth, with a track record of impactful work in ML systems, RL, or large-scale model training.

About Together AI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.