Member of Technical Staff - RL Inference
$180k - $440k • Palo Alto, CA
Posted 5d ago
About the job
SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. The RL infrastructure team is looking for an engineer to help with low precision RL training and inference. This role involves designing and optimizing the inference stack for all shapes of RL workloads, analyzing and addressing performance bottlenecks in large-scale RL systems, and working closely with the modeling team to efficiently implement novel RL techniques and algorithms.
Responsibilities
- Design and optimize the inference stack for RL workloads, from small scale ablations to production training runs.
- Analyze, profile, and address performance bottlenecks in large-scale RL systems.
- Collaborate with the modeling team to efficiently implement novel RL techniques and algorithms.
Requirements
- Experience building, debugging, and optimizing the efficiency of large-scale distributed systems.
- Experience with LLM inference.
- Proficiency in programming languages such as Python, C++, and/or Rust.
- Proficiency with frameworks such as PyTorch, Jax, and CUDA.
- Willingness to dive deep and solve hardcore problems at all levels of the stack.
Benefits
- Equity
- Comprehensive medical, vision, and dental coverage
- Access to a 401(k) retirement plan
- Short & long-term disability insurance
- Life insurance
- Various other discounts and perks