Staff Platform Engineer, Voice AI
$220k - $280k • Remote • San Francisco
Posted 1mo ago
Job Location
San Francisco
Tech Stack
Remote Work Policy
Fully remote
Categories
AI Infrastructure Engineer
About the job
Together AI is seeking a Staff Platform Engineer to lead the architecture of their Voice AI platform, which powers real-time voice agents at scale. This role involves setting the technical direction for how developers interact with the platform, from API primitives to autoscaling systems and multi-provider abstractions. The focus is on building robust, low-latency infrastructure for voice applications, which presents unique challenges compared to text inference, such as handling bidirectional audio streams and stateful connections. This is a foundational position on a small team, where decisions will shape the platform's architecture for years to come.
Responsibilities
- Own the architecture and reliability of the real-time API layer, setting technical direction for WebSocket and HTTP streaming APIs for STT and TTS at scale.
- Establish the reliability bar for production voice agents, including connection lifecycle, backpressure, graceful degradation, and reconnection.
- Lead autoscaling architecture for latency-sensitive voice workloads, designing and shipping orchestration systems for high-volume, real-time traffic.
- Define the voice API feature surface, making architectural decisions on word-level alignment, speaker diarization, audio format support, and multi-context WebSocket.
- Build the observability platform for voice infrastructure, designing latency breakdown pipelines, audio quality signal collection, and customer-facing dashboards.
- Own the multi-provider abstraction layer, architecting a normalization layer for consistent, provider-agnostic API behavior across model partners.
- Drive the interface between API and ML serving, defining the contract with the model serving stack to impact end-to-end latency and reliability.
- Raise the bar for developer experience across the platform through API design, documentation strategy, and integration patterns.
- Architect for future product surfaces, building systems that can serve as the foundation for new voice products.
Requirements
- 8+ years of experience building large-scale, real-time distributed systems with proven ownership of production systems.
- Deep expertise in real-time streaming infrastructure, including WebSocket server architecture, SSE, bidirectional streaming, connection multiplexing, and stateful protocol design.
- Expert-level TypeScript and Python proficiency, with strong systems-level thinking.
- Rust experience is a significant advantage.
- Senior distributed systems judgment in load balancing, autoscaling, rate limiting, and traffic shaping for latency-sensitive workloads.
- Deep Kubernetes expertise, including custom autoscalers, resource management, and health checking for stateful, streaming services.
- Strong technical leadership skills, including setting direction, influencing across teams, and bringing clarity to ambiguous problems.
- Sharp product intuition for developer platforms, with a focus on API ergonomics and developer experience.
- Proven ability to operate with autonomy on high-ambiguity, high-stakes problems, defining the right problem before optimizing the solution.
- Experience with audio and media protocols (WebRTC, g711, PCM encoding) is strongly preferred.
- Familiarity with ML model serving infrastructure and inference engines is a significant advantage.
- Full-stack experience (React, Next.js) for developer-facing tooling contributions is a plus.
- Bachelor's or Master's in Computer Science, Computer Engineering, or related field, or equivalent demonstrated experience.