Member of Technical Staff - ML Performance
New York • FullTime
Posted 4mo ago
Job Location
New York
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Machine Learning Engineer
About the job
AI needs a new infrastructure layer, and we're building it at Modal. We are looking for strong engineers with experience in making ML systems performant at scale. If you are interested in contributing to open-source projects and Modal’s container runtime to push language and diffusion models towards higher throughput and lower latency, we’d love to hear from you!
Responsibilities
- Contribute to open-source projects.
- Improve Modal's container runtime for higher throughput and lower latency for language and diffusion models.
Requirements
- 5+ years of experience writing high-quality, high-performance code.
- Experience working with torch, high-level ML frameworks, and inference engines (vLLM or TensorRT).
- Familiarity with Nvidia GPU architecture and CUDA.
- Experience with ML performance engineering, including debugging SM occupancy issues, rewriting algorithms for compute-bound performance, or eliminating host overhead.
- Familiarity with low-level operating system foundations (Linux kernel, file systems, containers) is a plus.