Member of Technical Staff - ML Performance

New York FullTime

Posted 4mo ago

Job Location

New York

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Machine Learning Engineer

About the job

AI needs a new infrastructure layer, and we're building it at Modal. We are looking for strong engineers with experience in making ML systems performant at scale. If you are interested in contributing to open-source projects and Modal’s container runtime to push language and diffusion models towards higher throughput and lower latency, we’d love to hear from you!

Responsibilities

  • Contribute to open-source projects.
  • Improve Modal's container runtime for higher throughput and lower latency for language and diffusion models.

Requirements

  • 5+ years of experience writing high-quality, high-performance code.
  • Experience working with torch, high-level ML frameworks, and inference engines (vLLM or TensorRT).
  • Familiarity with Nvidia GPU architecture and CUDA.
  • Experience with ML performance engineering, including debugging SM occupancy issues, rewriting algorithms for compute-bound performance, or eliminating host overhead.
  • Familiarity with low-level operating system foundations (Linux kernel, file systems, containers) is a plus.

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.