Senior Backend Engineer, Inference Platform

$160k - $250k Remote San Francisco

Posted 1mo ago

Remote Work Policy

Fully remote

Categories

LLM Engineer

About the job

Together AI is building the Inference Platform to bring advanced generative AI models to the world, powering multi-tenant serverless workloads and dedicated endpoints. This role offers a unique opportunity to optimize latency and fully utilize tens of thousands of GPUs, working hands-on with cutting-edge hardware. You will collaborate directly with research teams to productionize frontier models and engage with the open-source community, contributing to projects that push the boundaries of inference performance and efficiency.

Responsibilities

  • Build and optimize global and local request routing for low-latency load balancing.
  • Develop auto-scaling systems for dynamic resource allocation and meeting SLOs.
  • Design systems for multi-tenant traffic shaping, including resource allocation and request handling.
  • Engineer trade-offs between latency and throughput for efficient workload serving.
  • Optimize prefix caching to reduce model compute and speed up responses.
  • Collaborate with ML researchers to productionize new model architectures at scale.
  • Continuously profile and analyze system performance to identify and implement optimizations.

Requirements

  • 5+ years of experience building large-scale, fault-tolerant, distributed systems and API microservices.
  • Strong background in designing, analyzing, and improving efficiency, scalability, and stability of complex systems.
  • Excellent understanding of low-level OS concepts: multi-threading, memory management, networking, and storage performance.
  • Expert-level programming in Rust, Go, Python, or TypeScript.
  • Knowledge of modern LLMs and generative models and their production serving is a plus.
  • Experience with the open-source inference ecosystem (e.g., SGLang, vLLM, NVIDIA Dynamo) is highly valuable.
  • Experience with Kubernetes or container orchestration is a strong plus.
  • Familiarity with GPU software stacks (CUDA, Triton, NCCL) and HPC technologies (InfiniBand, NVLink, MPI) is a plus.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience.

Benefits

  • Competitive compensation
  • Startup equity
  • Health insurance
  • Other competitive benefits

About Together AI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.