Staff Research Engineer - Multimodal Generative Modelling

Remote Europe FullTime

Posted 1mo ago

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

AI Research Engineer

About the job

Synthesia is seeking a Staff Research Engineer to join their Voice team and contribute to the company's long-term vision of building the best human-interactive models. These advanced systems will go beyond simple conversation to perceive, respond to, and react to user actions and emotions in real-time. The role involves defining and driving this broader vision across teams, proposing ambitious research directions, and taking ownership of critical component design and implementation. You will collaborate closely with the voice lead, video teams, and other senior members to develop models that combine text, audio, and video for a seamless, interactive experience, moving beyond traditional turn-taking in speech models.

Responsibilities

  • Shape the roadmap for new model capabilities and customer-facing functionality.
  • Propose novel multi-modal system architectures, particularly for text and voice.
  • Develop and evaluate streaming and conversational systems for low-latency, interactive voice-video synthesis.
  • Design solutions that enhance emotional expressiveness and natural interaction.
  • Implement and bring designs to life, from pretraining through post-training.
  • Integrate and test novel architectures like neural codecs, diffusion, and flow-matching.
  • Define new evaluation metrics for conversational systems, including latency-aware and interaction-based measurements.
  • Track the latest research in audio-visual diffusion, autoregressive models, neural codecs, and multimodal LLMs.
  • Curate new datasets to complement existing data.
  • Lead post-training initiatives such as DPO, fine-tuning, and distillation.
  • Ship models to production with optimized runtime and address customer feedback.

Requirements

  • Ability to bring novel ideas and designs that advance the field of interactive multimodal systems.
  • Strong understanding of generative modelling, ideally applied to sequential or multimodal data.
  • Hands-on experience with large language models or similar transformer-based architectures.
  • High proficiency in PyTorch, including distributed training and model optimization.
  • Solid grasp of time-series modeling and tokenization, preferably in audio, speech, or video contexts.
  • Demonstrated ability to prototype quickly, test hypotheses, and iterate efficiently.
  • Proven experience training deep learning models end-to-end.
  • Strong general software engineering skills for contributions to shared research infrastructure.
  • Experience shipping a generative model into a live product at meaningful scale.
  • Experience working on conversational or interactive systems where latency, responsiveness, and user experience were key constraints.
  • Experience working on LLMs with large-scale trainings leading to models with decent reasoning capabilities.
  • Experience owning a research problem end-to-end, from architecture proposal to production deployment.
  • Experience collaborating across modalities or teams to ship a unified system.

About synthesia.io

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.