Senior Platform Engineer, Voice AI
$200k - $260k • Remote • San Francisco
Posted 1mo ago
Job Location
San Francisco
Tech Stack
Remote Work Policy
Fully remote
Categories
Applied AI Engineer
About the job
Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications with best-in-class latency and reliability. We are seeking a Senior Platform Engineer to take ownership of the API and infrastructure layer for voice workloads. You will develop the real-time WebSocket and HTTP APIs used by developers to deploy voice experiences, design autoscaling for latency-sensitive streaming workloads, and ensure the reliability of our multi-provider voice platform for production voice agents handling millions of calls. This is a critical, foundational role on a small, high-impact team, defining how developers interact with our voice platform as we scale.
Responsibilities
- Build and harden real-time WebSocket and HTTP streaming APIs for STT and TTS, managing connection lifecycles, backpressure, error handling, and reconnection for production voice agents.
- Design and implement autoscaling for voice model endpoints to handle bursty, real-time traffic, considering concurrent connection limits, streaming state, and latency ceilings.
- Implement voice-specific API features such as word-level alignment, real-time speaker diarization, flexible audio format support, pronunciation controls, and multi-context WebSocket support.
- Develop voice-specific observability tools, including latency breakdowns, audio quality signals, and dashboards for debugging.
- Manage multi-model normalization across various model partners to ensure consistent API behavior.
- Collaborate with ML engineering on the interface between the API layer and model serving stack to meet end-to-end latency and reliability requirements.
- Contribute to developer experience through API design, documentation, integration cookbooks, and showcasing best practices for voice agents.
- Lay the groundwork for future product development.
Requirements
- 5+ years of experience building large-scale, real-time distributed systems and API services.
- Deep expertise in real-time streaming infrastructure, including WebSocket server architecture, Server-Sent Events, bidirectional streaming, connection multiplexing, and stateful protocol design.
- Expert-level proficiency in TypeScript and Python; Rust experience is a plus.
- Strong distributed systems fundamentals: load balancing, autoscaling, rate limiting, and traffic shaping for latency-sensitive workloads.
- Experience with Kubernetes, including custom autoscalers, resource management, and health checking for stateful services.
- Strong product sense with a focus on API ergonomics and developer needs for voice applications.
- Comfort working in a fast-paced, small team environment where wearing multiple hats is expected.
- Experience with audio or media protocols (WebRTC, g711, PCM encoding) is a strong plus.
- Familiarity with ML model serving infrastructure and inference engines is a plus.
- Full-stack experience (React, Next.js) is a nice-to-have.
- Bachelor's or Master's degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience.
Benefits
- Competitive compensation
- Startup equity
- Health insurance
- Other competitive benefits