Member of Technical Staff - Mid-Training Infra
San Francisco, CA • FullTime
Posted 4mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
Reflection is a research lab dedicated to making intelligence open and accessible. We build open models that empower individuals to control their intelligence and shape the future of AI. As a Member of Technical Staff focused on Mid-Training Infrastructure, you will be instrumental in designing, building, and operating large-scale GPU infrastructure crucial for high-throughput model inference and mid-training workloads. This role involves developing systems that support synthetic data generation and reinforcement learning pipelines at scale, as well as building high-performance inference platforms capable of serving and evaluating models across thousands of GPUs.
Responsibilities
- Design, build, and operate large-scale GPU infrastructure for high-throughput model inference and mid-training workloads.
- Develop systems powering synthetic data generation and reinforcement learning pipelines at scale.
- Build high-performance inference platforms for serving and evaluating models across thousands of GPUs.
- Optimize throughput, latency, and GPU utilization for large language model inference and rollout workloads.
- Build infrastructure supporting reinforcement learning pipelines, including large-scale rollout generation, evaluation, and policy improvement loops.
- Support research teams with distributed RL workloads and large-scale model evaluation infrastructure.
- Improve model execution performance through kernel-level optimization, model parallelism, and GPU runtime enhancements.
- Develop distributed systems for large-scale synthetic data generation and RL-driven training workflows.
- Diagnose and resolve performance bottlenecks across inference runtimes, GPU kernels, networking, and distributed compute systems.
Requirements
- Experience deploying and operating large-scale GPU systems for inference or model serving.
- Several years of hands-on experience building and running production infrastructure.
- Strong understanding of GPU performance characteristics and optimization techniques.
- Experience with modern inference frameworks like SGLang, Megatron, or similar high-performance LLM runtimes.
- Familiarity with distributed reinforcement learning infrastructure or rollout generation systems.
- Experience optimizing throughput for large-scale model execution workloads.
- Experience working with GPU kernels or low-level performance optimization.
- Familiarity with infrastructure for synthetic data pipelines or RL training workflows.
- Experience debugging performance issues across GPU, networking, and distributed execution layers.
Benefits
- Top-tier compensation: Salary and equity structured to recognize and retain talent globally.
- Stock options
- Comprehensive medical, dental, vision, and life insurance
- Annual wellness allowance
- Provided lunch and dinner in the office daily
- 22 weeks paid parental leave for all new parents
- Unlimited paid time off in the U.S.
- 30 days paid time off in the U.K.
- Sponsorship support for visas and long-term immigration pathways
- Regular off-sites, happy hours, and team celebrations