Member of Technical Staff - Research, Inference
New York • FullTime
Posted 2mo ago
Job Location
New York
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Research Engineer
About the job
AI requires a new infrastructure layer, and Modal is building it. We provide customers with instant GPU access, sub-second container starts, and native storage, simplifying the deployment and serving of low-latency inference, model fine-tuning, and production-ready sandboxes at scale. We are seeking a Member of Technical Staff focused on Research and Inference to join our team. This role involves hands-on inference research, identifying high-impact areas, and driving them to completion. The primary focus will be on improving cost per token and tail latency for customer workloads.
Responsibilities
- Own end-to-end inference research projects, including speculative decoding, disaggregated prefill/decode, quantization (FP8, INT4), KV-cache and memory management, and autoscaling for serverless traffic.
- Train custom speculators using real production traffic and integrate learnings back into target models.
- Collaborate directly with customers to deploy and tune models, feeding insights back into research.
- Expand collaborations with external research labs on projects like DFlash, specdec, multimodal inference, and Flash Attention 4 kernels.
- Partner with engineering to translate cutting-edge serving techniques into products, such as primitives for disaggregation, fast weight refresh, and production observability.
- Contribute to shaping the future research agenda based on project outcomes.
Requirements
- Research-leaning or systems background in LLM inference with demonstrable work.
- Proficiency in the LLM serving stack, from kernels and quantization to schedulers and autoscaling.
- A history of shipping impactful research or systems used by others.
- Ability to independently drive research projects from conception to completion.
- Willingness to work collaboratively and openly with the team.
- Ability to work in-person in our NYC or San Francisco office.