Senior Machine Learning Engineer

Hybrid

Posted 10d ago

Remote Work Policy

On-site

Categories

Machine Learning Engineer

About the job

You will help define how machine learning models run across Cloudflare’s global network, from frontier open LLMs and real-time voice models to customer-deployed models served on heterogeneous GPUs and next-generation accelerators. You’ll work with systems engineers, product teams, hardware partners, and AI/ML engineers to bring models into production with low latency, strong reliability, and efficient resource use. This role combines applied ML, inference optimization, evaluation, and production engineering, with a focus on benchmarking models, improving serving performance, validating quality, and building tooling that helps Cloudflare and its customers ship AI applications at Internet scale.

Responsibilities

  • Develop, optimize, and productionize machine learning models for Cloudflare’s serverless inference platform, focusing on performance, reliability, and model quality.
  • Build benchmarking and evaluation frameworks to measure latency, throughput, cost efficiency, and model behavior across various model families.
  • Improve inference performance through techniques like quantization, batching, caching, model compilation, runtime tuning, and accelerator-aware optimization.
  • Partner with systems engineers to integrate models into Cloudflare’s distributed inference infrastructure across a heterogeneous fleet of GPUs and accelerators.
  • Drive improvements to model deployment workflows, including validation, rollout safety, observability, regression testing, and operational readiness.
  • Collaborate with product and engineering teams to translate customer requirements into scalable ML capabilities for Workers AI.
  • Mentor engineers, contribute to technical direction, and raise the quality bar for production ML engineering practices.

Requirements

  • Experience building, optimizing, and operating machine learning models in production environments.
  • Strong proficiency with Python and modern ML frameworks such as PyTorch, TensorFlow, JAX, or equivalent.
  • Hands-on experience with inference optimization techniques for large-scale models.
  • Experience with large-scale inference serving frameworks or runtimes.
  • Familiarity with LLMs, speech models, vision models, embeddings, multimodal models, retrieval-augmented generation, or other modern deep learning architectures.
  • Experience optimizing models for GPUs or specialized accelerators.
  • Strong understanding of production ML concerns, including evaluation, monitoring, model regressions, rollout safety, and reliability.
  • Ability to work across ML and systems boundaries, including familiarity with distributed systems, networking, or serverless platforms.
  • Track record of leading complex technical projects and mentoring other engineers.

About Cloudflare

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.