Staff / Principal Machine Learning Engineer, Serving - Switzerland

Remote Switzerland FullTime

Posted 5mo ago

Job Location

Switzerland

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Machine Learning Engineer

About the job

Inworld is seeking a Staff/Principal Machine Learning Engineer to join their research lab focused on building top-ranked real-time voice models. These models power large consumer-facing AI applications across various sectors. The role involves optimizing real-time inference, developing best-in-class APIs and products, and contributing to the research and development of state-of-the-art models. The ideal candidate is a fast learner who thrives in ambiguity and can demonstrate a strong portfolio of built, broken, and understood systems. This position emphasizes impact, shipping stable code, and a deep understanding of the underlying logic behind engineering decisions.

Responsibilities

  • Optimize inference for state-of-the-art models using modern serving frameworks.
  • Accelerate model performance through techniques like quantization, distillation, caching, continuous batching, paged attention, and speculative decoding.
  • Develop and optimize high-performance systems using C++, CUDA, Rust, or highly optimized Python.
  • Profile code and maximize performance on NVIDIA GPUs.
  • Implement distributed systems and scaling solutions using Kubernetes, Ray, and custom load balancing.
  • Handle thousands of concurrent connections reliably.
  • Take ownership of models from research to production, including containerization and serving optimization.
  • Collaborate daily with US-based leadership and engineering teams.

Requirements

  • Deep understanding of modern serving frameworks and techniques (e.g., vLLM, TRT-LLM).
  • Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding.
  • Proficiency in C++, CUDA, Rust, or highly optimized Python.
  • Experience with profiling code and optimizing for NVIDIA GPUs.
  • Experience with Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference, and handling thousands of concurrent connections.
  • Demonstrated experience with non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups.
  • Ability to take a model from research to production, ensuring reliable serving.
  • PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems.
  • Professional fluency in English (written and spoken).

About inworld

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.