ML Research Engineer, ML Systems
$190k - $237k • San Francisco, CA; Seattle, WA; New York, NY
Posted 2mo ago
Job Location
San Francisco, CA; Seattle, WA; New York, NY
Tech Stack
Remote Work Policy
On-site
Categories
Machine Learning Engineer
About the job
Scale's ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. This platform powers MLEs, researchers, data scientists, and operators for fast and automatic training and evaluation of LLMs, as well as data quality evaluation. You will work closely across Scale’s ML teams and researchers to build the foundation platform that supports all our ML research and development, optimizing it to enable the next generation of LLM training, inference, and data curation. If you are excited about shaping the future of AI via fundamental innovations, we would love to hear from you!
Responsibilities
- Build, profile, and optimize the training and inference framework.
- Collaborate with ML teams to accelerate their research and development.
- Enable ML teams to develop the next generation of models and data curation.
- Research and integrate state-of-the-art technologies to optimize the ML system.
Requirements
- Strong excitement about system optimization.
- Experience with multi-node LLM training and inference.
- Experience developing large-scale distributed ML systems.
- Strong software engineering skills.
- Proficiency in frameworks and tools such as CUDA, Pytorch, transformers, flash attention.
- Strong written and verbal communication skills.
- Ability to operate in a cross-functional team environment.
Benefits
- Base salary
- Equity
- Comprehensive health, dental, and vision coverage
- Retirement benefits
- Learning and development stipend
- Generous PTO
- Commuter stipend