Member of Technical Staff - Research, Inference

New York FullTime

Posted 2mo ago

Job Location

New York

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Research Engineer

About the job

AI requires a new infrastructure layer, and Modal is building it. We provide customers with instant GPU access, sub-second container starts, and native storage, simplifying the deployment and serving of low-latency inference, model fine-tuning, and production-ready sandboxes at scale. We are seeking a Member of Technical Staff focused on Research and Inference to join our team. This role involves hands-on inference research, identifying high-impact areas, and driving them to completion. The primary focus will be on improving cost per token and tail latency for customer workloads.

Responsibilities

  • Own end-to-end inference research projects, including speculative decoding, disaggregated prefill/decode, quantization (FP8, INT4), KV-cache and memory management, and autoscaling for serverless traffic.
  • Train custom speculators using real production traffic and integrate learnings back into target models.
  • Collaborate directly with customers to deploy and tune models, feeding insights back into research.
  • Expand collaborations with external research labs on projects like DFlash, specdec, multimodal inference, and Flash Attention 4 kernels.
  • Partner with engineering to translate cutting-edge serving techniques into products, such as primitives for disaggregation, fast weight refresh, and production observability.
  • Contribute to shaping the future research agenda based on project outcomes.

Requirements

  • Research-leaning or systems background in LLM inference with demonstrable work.
  • Proficiency in the LLM serving stack, from kernels and quantization to schedulers and autoscaling.
  • A history of shipping impactful research or systems used by others.
  • Ability to independently drive research projects from conception to completion.
  • Willingness to work collaboratively and openly with the team.
  • Ability to work in-person in our NYC or San Francisco office.

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.