Principal ML Platform Engineer

Remote Europe FullTime

Posted 5mo ago

Job Location

Europe

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

Synthesia is seeking a Principal Engineer to join their ML Platform team. This role focuses on building and operating the systems that enable researchers and product teams to train, serve, and deploy generative models efficiently and reliably. The team's work encompasses research infrastructure, production serving systems, internal tooling, and platform interfaces, with a growing emphasis on automation and agent-oriented workflows. This is a hands-on individual contributor position with significant ownership, where you will influence the evolution of the ML platform as it scales.

Responsibilities

  • Design and improve platform systems for model training, evaluation, and production serving.
  • Build infrastructure and tooling to enhance the reliability, scalability, and cost-efficiency of ML workloads.
  • Develop internal tools and workflows that are easily operated by both humans and agents.
  • Work on the architecture for deploying, serving, and operating models across research and product environments.
  • Improve the scheduling, monitoring, and debugging of workloads on GPUs and cloud infrastructure.
  • Develop internal tools, abstractions, and agentic systems to reduce operational overhead for researchers and engineers.
  • Drive improvements in observability, automation, reliability, and developer experience.
  • Collaborate with researchers and product engineers to address pain points and build robust platform capabilities.
  • Contribute to technical direction and make pragmatic architectural tradeoffs for platform growth.

Requirements

  • Strong experience building or operating production systems with a focus on reliability, scalability, and maintainability.
  • A systems mindset, considering bottlenecks, failure modes, interfaces, resource usage, and long-term operability.
  • Solid hands-on experience with cloud infrastructure, Linux, and infrastructure automation.
  • Experience with Kubernetes and operating distributed workloads in production.
  • Strong coding skills, ideally in Python or similar backend/tooling languages.
  • Strong judgment regarding automation leverage versus human control and reliability.
  • Experience building internal platforms, developer tooling, or infrastructure abstractions for other engineers.
  • Comfort working in ambiguous environments and taking ownership of open-ended technical problems.
  • A pragmatic approach focused on solving the right problem well, avoiding over-engineering.

About synthesia.io

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.