Senior Machine Learning Engineer, Voice AI

$200k - $260k Remote San Francisco

Posted 1mo ago

Remote Work Policy

Fully remote

Categories

Machine Learning Engineer

About the job

Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Senior ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro to achieve frontier-level latency and throughput. You will focus on unique voice inference challenges such as streaming audio, tokenization, and real-time latency budgets, shaping how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.

Responsibilities

  • Own the model serving stack for Together's voice platform (STT, TTS, speech-to-speech).
  • Optimize inference performance for voice models, targeting best-in-class TTFB, throughput, and GPU utilization.
  • Productionize voice models on serverless and dedicated endpoints, including batching, streaming inference, and memory management for audio workloads.
  • Build and maintain a voice model evaluation framework to measure WER, naturalness, latency, and pronunciation accuracy.
  • Enable new model architectures in the serving stack, including audio-native LLMs and speech-to-speech systems.
  • Collaborate with model partners to integrate and optimize their models on Together's infrastructure.
  • Profile and debug performance across the full inference stack, shipping measurable improvements.
  • Ensure the serving layer meets latency and reliability requirements for real-time voice APIs.
  • Contribute to voice model fine-tuning capabilities for STT and TTS.
  • Lay the groundwork for multiple new products.

Requirements

  • 5+ years of experience in ML engineering, focusing on model serving, inference optimization, or ML infrastructure.
  • Hands-on experience with LLM serving engines (vLLM, SGLang, TensorRT-LLM, or similar), including modifying engine internals.
  • Strong proficiency in Python and PyTorch.
  • Experience with GPU profiling and optimization (CUDA, memory management, kernel-level debugging).
  • Track record of shipping ML systems to production with measurable performance improvements.
  • Strong product sense, understanding the needs of developers building voice apps.
  • Comfort working on a small, early-stage team with a fast-paced environment.
  • Experience with speech and audio ML (ASR, TTS architectures, audio signal processing) is a strong plus.
  • Familiarity with audio codecs and tokenization schemes (SNAC, Encodec, DAC) is a plus.
  • Experience training or fine-tuning speech models is a plus.
  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field, or equivalent practical experience.

Benefits

  • Competitive compensation
  • Startup equity
  • Health insurance
  • Other competitive benefits

About Together AI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.