Member of Technical Staff (AI Inference Engineer)

London FullTime

Posted 5mo ago

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

We are seeking an AI Inference Engineer to join our dynamic team. This role is central to Perplexity's operations, as you will build and manage the inference engine that powers every query. You will deploy a variety of model architectures at scale, focusing on meeting stringent latency and cost requirements. Our technology stack includes Rust, Python, CUDA, and the CuTe DSL.

Responsibilities

  • Support transformer-based retrieval, text-generation, and multimodal models within our inference infrastructure, covering aspects from weight loading and request scheduling to KV-cache management and API Gateway integration.
  • Migrate in-house CUDA kernels to NVIDIA's CuTe DSL to ensure compatibility with current hardware (GB200) and future systems.
  • Develop our internal Rust-based inference server to address Python-related challenges and scale with increasing traffic.
  • Profile and optimize performance bottlenecks across the entire stack, from network ingress to continuous batching and GPU kernel interleaving.
  • Build and maintain dashboards, alerts, and automated remediation systems to proactively identify and resolve regressions, and respond to production incidents.

Requirements

  • Deep experience with GPU programming and performance optimization (CUDA, Triton, CUTLASS, or similar).
  • Understanding of modern LLM architectures and experience deploying them reliably in production.
  • Proven experience building and operating production distributed systems under significant load, particularly performance-critical ones.
  • Proficiency in working across multiple languages and layers, including Rust for serving, Python for model code, and CUDA/CuTeDSL for kernels.
  • Ability to own problems end-to-end, from understanding research papers to debugging production issues.
  • Self-directed and capable of thriving in fast-moving environments with evolving priorities.
  • 3+ years of professional software engineering experience with a focus on ML inference or high-performance systems.
  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
  • Knowledge of common LLM architectures and inference optimization techniques (e.g., quantization, speculative decoding).

Benefits

  • Equity as part of the total compensation package.

About Perplexity AI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.