Member of Technical Staff - Web Crawl Engineer
San Francisco, CA • FullTime
Posted 1mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Reflection is a research lab dedicated to making intelligence open and accessible. As a Member of Technical Staff on the Data Team, you will be instrumental in building and operating large-scale web crawling systems. These systems are crucial for discovering, acquiring, and processing internet content, directly influencing the capabilities of frontier AI systems. You will own the infrastructure for web-scale data collection, from initial URL discovery to distributed crawling and dataset delivery, working closely with researchers to identify and acquire high-value content efficiently. This role is perfect for engineers passionate about distributed systems, large-scale crawler optimization, and tackling the unique challenges of modern web data collection.
Responsibilities
- Build and operate web-scale crawling infrastructure for continuous data collection across billions of URLs.
- Design and optimize systems for URL discovery, prioritization, scheduling, and crawl orchestration.
- Develop distributed crawlers for efficient content acquisition while respecting site constraints.
- Build systems for content extraction, rendering, parsing, and normalization across diverse web formats.
- Improve crawl coverage, freshness, efficiency, and quality through measurement and experimentation.
- Design infrastructure for large-scale recrawling, change detection, and incremental updates.
- Develop specialized crawlers for high-value domains, dynamic websites, and difficult-to-access content.
- Analyze crawl performance and web coverage to identify gaps and inefficiencies.
- Build observability, monitoring, and reliability systems for large-scale crawl operations.
- Debug production issues and enhance the performance, scalability, and resilience of crawling infrastructure.
Requirements
- Experience building large-scale web crawling, search indexing, content acquisition, or internet-scale data collection systems.
- Strong understanding of crawling architectures, URL frontier management, scheduling, and distributed crawl coordination.
- Experience with large-scale distributed systems using frameworks like Ray, Spark, Beam, or Flink.
- Familiarity with content extraction, HTML parsing, browser automation, rendering systems, and modern web technologies.
- Experience operating systems that process petabyte-scale datasets.
- Strong systems engineering skills, including reliability, observability, performance optimization, and debugging.
- Experience designing experiments and using data to improve crawl quality, coverage, and efficiency.
- Excellent communication skills and ability to reason about system tradeoffs and operational constraints.
- Passion for web-scale systems and internet information collection.
- Curiosity about web data's influence on model capabilities and willingness to iterate based on results.
- Comfort balancing crawl quality, coverage, freshness, and operational efficiency.
- Enjoyment working at the intersection of distributed systems, data infrastructure, and AI.
Benefits
- Top-tier compensation (salary and equity)
- Stock options
- Comprehensive medical, dental, vision, and life insurance
- Annual wellness allowance
- Provided lunch and dinner in the office
- 22 weeks paid parental leave
- Unlimited paid time off (US) / 30 days paid time off (UK)
- Visa sponsorship support
- Regular off-sites, happy hours, and team celebrations