Staff Machine Learning Engineer, Voice AI
$220k - $280k • Remote • San Francisco
Posted 1mo ago
Job Location
San Francisco
Tech Stack
Remote Work Policy
Fully remote
Categories
Machine Learning Engineer
About the job
Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Staff ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro, focusing on pushing latency and throughput boundaries. You will address unique challenges in voice inference, such as streaming audio and real-time latency, and shape the future of how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.
Responsibilities
- Own the voice inference roadmap, defining and executing the technical strategy for optimizing STT, TTS, and speech-to-speech models.
- Architect and implement systems for best-in-class inference performance, targeting leading TTFB, throughput, and GPU utilization for voice workloads.
- Design the serving architecture for serverless and dedicated endpoints, including batching strategies, streaming inference pipelines, and memory management for real-time audio.
- Build a rigorous and extensible evaluation framework for STT and TTS models, establishing internal benchmarks.
- Anticipate and enable emerging model paradigms like audio-native LLMs and end-to-end speech-to-speech systems.
- Lead deep collaboration with model partners for integration, optimization, and performance accountability.
- Conduct systematic profiling and root-cause analysis to diagnose and resolve performance problems in the stack.
- Influence platform architecture to meet the latency and reliability demands of real-time voice APIs.
- Lead the technical direction for enabling customer fine-tuning of STT and TTS models.
- Architect systems that support multiple new voice products with minimal rework, focusing on platform-level solutions.
Requirements
- 8+ years of ML engineering experience with a focus on model serving, inference optimization, or ML infrastructure at production scale.
- Deep, practical expertise in LLM serving engines (e.g., vLLM, SGLang, TensorRT-LLM), including modifying internals and debugging edge cases.
- Expert-level Python and PyTorch proficiency, with strong command of GPU optimization (CUDA kernels, memory hierarchies, profiling toolchains).
- Proven system design judgment with a track record of making architectural decisions that held up at scale.
- Strong technical leadership, operating with high autonomy and raising engineering quality standards.
- Sharp product intuition for developer tooling, understanding the needs of voice application developers.
- Proven ability to move fast in ambiguous environments on early-stage or platform teams.
- Strong foundation in speech and audio ML (ASR/TTS architectures, audio signal processing) is strongly preferred.