Full-Stack Software Engineer, Reinforcement Learning
San Francisco, CA | New York City, NY
Posted 16d ago
Job Location
San Francisco, CA | New York City, NY
Tech Stack
Remote Work Policy
On-site
Categories
Applied AI Engineer
About the job
As a Full-Stack Software Engineer in Reinforcement Learning (RL), you will be instrumental in building the platforms, tools, and interfaces essential for environment creation, data collection, and training observability. Your work will directly impact the quality of data used to train Anthropic's next-generation AI models. You will own product surfaces from end-to-end, encompassing backend services, APIs, and web UIs used by researchers, external vendors, and data labelers. The role emphasizes shipping polished, reliable products quickly, even when faced with ambiguous, high-stakes problems. This team operates at a rapid pace, focusing on judgment and taste to meet researcher needs, iterating on data collection strategies to distill expert knowledge into models within short feedback loops.
Responsibilities
- Build and extend web platforms for RL environment creation, management, and quality review.
- Develop vendor-facing interfaces and tooling for training environment iteration.
- Design and implement platforms for large-scale human data collection, including labeling workflows and quality assurance.
- Build evaluation dashboards and observability UIs for real-time insights into training runs.
- Create backend services and APIs connecting environment authoring, data collection, and training infrastructure.
- Build and expand scalable code data generation pipelines.
- Develop onboarding automation and documentation tooling for rapid user ramp-up.
- Partner with researchers and operations teams to translate requirements into well-designed products.
Requirements
- Strong software engineering fundamentals with full-stack experience, from database schema to frontend.
- Proficiency in Python and a modern web stack (e.g., React, TypeScript).
- Proven track record of shipping systems that solved hard problems and improved team efficiency.
- High agency: ability to identify needs and drive projects forward independently.
- Experience building intuitive interfaces for both technical and non-technical users.
- Clear communication skills with researchers, operations teams, and engineers.
- Ability to thrive in a fast-moving environment with shifting priorities.
- Commitment to Anthropic's mission of building safe, beneficial AI.