Data Infrastructure Engineer, Pre-training
San Francisco, CA
Posted 10d ago
About the job
Anthropic is seeking a Staff level Engineer to join our Pre-training team, responsible for developing the next generation of large language models. In this role, you will work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems. Our mission is to ensure that transformative AI systems are aligned with human interests.
Responsibilities
- Design and implement highly performant, reproducible, and traceable data processing infrastructure for large language model training.
- Develop and maintain core processing primitives like tokenization, deduplication, and chunking with a focus on scalability.
- Build robust systems for data quality assurance and validation at scale.
- Collaborate with research teams to implement novel data processing architectures.
- Build and operate end-to-end data pipelines that transform raw web-scale corpora into training-ready datasets.
Requirements
- 5+ years of experience outside of internships.
- Strong software engineering skills with experience building high-throughput, fault-tolerant distributed systems.
- Hands-on experience with distributed computing frameworks, particularly Apache Spark.
- Excellent problem-solving skills and attention to detail.
- Strong communication skills and ability to work in a collaborative environment.
- Advanced degree in Computer Science or related field.
- Experience with language model training infrastructure.
- Background in Data Infrastructure, MLOps, or ML infrastructure.
Benefits
- Annual compensation range: $500,000 - $850,000 USD
- Visa sponsorship available