Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI
$265k - $331k • San Francisco, CA; New York, NY
Posted 2mo ago
Job Location
San Francisco, CA; New York, NY
Tech Stack
Remote Work Policy
On-site
Categories
Machine Learning Engineer
About the job
Scale is seeking a Machine Learning Systems Research Engineer to join their Enterprise ML Research Lab. This role will focus on building algorithms for a next-generation Agent RL training platform, supporting large-scale training, and integrating state-of-the-art technologies to optimize ML systems. You will collaborate with other ML researchers and engineers who apply these algorithms to client use cases, including AI cybersecurity firewalls and healthtech search models. If you are passionate about shaping the future of AI, this is an exciting opportunity to contribute to cutting-edge advancements in enterprise GenAI.
Responsibilities
- Build, profile, and optimize the training and inference framework.
- Post-train state-of-the-art models for enterprise engagements.
- Collaborate with ML teams to accelerate research and development.
- Create next-generation agent training algorithms for multi-agent/multi-tool rollouts.
Requirements
- 1-3 years of LLM training experience in a production environment.
- Passion for system optimization.
- Experience with post-training methods like RLHF/RLVR and algorithms like PPO/GRPO.
- Knowledge of modern GPU cluster architecture.
- Experience with multi-node LLM training and inference.
- Strong software engineering skills with proficiency in CUDA, Pytorch, transformers, and flash attention.
- Strong written and verbal communication skills for cross-functional collaboration.
- PhD or Master's degree in Computer Science or a related field.
Benefits
- Base salary
- Equity
- Comprehensive health, dental, and vision coverage
- Retirement benefits
- Learning and development stipend
- Generous PTO
- Commuter stipend