Principal Research & Engineering, Realtime Voice AI
$400k - $550k • Palo Alto, California, United States
Posted 1mo ago
Remote Work Policy
On-site
Categories
AI Research Engineer
About the job
Inflection AI is seeking a hands-on technical leader to define and build the company's real-time Voice AI stack. This role involves shaping the future of spoken interactions with emotionally intelligent AI, focusing on speech models, streaming systems, voice-agent runtime, and evaluation. The ideal candidate will partner across research, engineering, product, and design to deliver responsive, trustworthy, and useful voice agents for enterprise settings.
Responsibilities
- Establish the technical roadmap for real-time Voice AI, including streaming ASR, TTS, speech-to-speech, speech LLMs, turn-taking, barge-in, latency, and reliability.
- Utilize a 1,000 GPU cluster for performance benchmarking and experimentation.
- Determine build-vs-buy-vs-train strategies for core audio, speech, and real-time interaction components.
- Direct research and engineering efforts on speech quality, naturalness, expressiveness, emotional fit, controllability, and production readiness.
- Collaborate with infrastructure, product, design, and agentic AI teams to deploy voice agents for enterprise workflows.
- Develop evaluation systems for voice quality beyond standard WER, measuring clarity, emotional appropriateness, interruption handling, task success, user preference, latency, and reliability.
- Refine production voice behavior by debugging across runtime, model, evaluation, data, and product layers.
- Mentor and coach a team specializing in speech research, audio infrastructure, real-time systems, and evaluation.
Requirements
- Experience leading or serving as a principal Research and Engineering contributor to real-time voice, speech, audio AI, or conversational AI systems in production.
- Experience with one or more of: streaming ASR, TTS, speech-to-speech systems, speech LLMs, audio tokenization, multimodal models, barge-in, low-latency inference, or real-time agents.
- Strong technical judgment across both speech modeling and production systems.
- Ability to define voice quality in terms of user and customer outcomes, not only offline model metrics.
- Experience designing or using evaluation systems that capture real user experience.
- Strong product intuition for natural, trustworthy, emotionally appropriate voice interactions.
- Ability to lead senior technical talent while staying close to the code, architecture, and debugging work.
- Bachelor’s degree or equivalent in a related field.
Benefits
- Meaningful equity component
- Robust medical, dental and vision options with employer contributions for HSA, FSA and DFSA
- 401k matching program
- Flexible Time Off
- 10 paid holidays
- 5 days sick leave
- Parental, Medical and Family care leave
- Generous cell-phone, wellness and office set up stipends
- Support of country-specific visa needs for international employees living in the Bay Area