Member of Technical Staff - Web Crawl Engineer

San Francisco, CA FullTime

Posted 1mo ago

Job Location

San Francisco, CA

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Reflection is a research lab dedicated to making intelligence open and accessible. As a Member of Technical Staff on the Data Team, you will be instrumental in building and operating large-scale web crawling systems. These systems are crucial for discovering, acquiring, and processing internet content, directly influencing the capabilities of frontier AI systems. You will own the infrastructure for web-scale data collection, from initial URL discovery to distributed crawling and dataset delivery, working closely with researchers to identify and acquire high-value content efficiently. This role is perfect for engineers passionate about distributed systems, large-scale crawler optimization, and tackling the unique challenges of modern web data collection.

Responsibilities

  • Build and operate web-scale crawling infrastructure for continuous data collection across billions of URLs.
  • Design and optimize systems for URL discovery, prioritization, scheduling, and crawl orchestration.
  • Develop distributed crawlers for efficient content acquisition while respecting site constraints.
  • Build systems for content extraction, rendering, parsing, and normalization across diverse web formats.
  • Improve crawl coverage, freshness, efficiency, and quality through measurement and experimentation.
  • Design infrastructure for large-scale recrawling, change detection, and incremental updates.
  • Develop specialized crawlers for high-value domains, dynamic websites, and difficult-to-access content.
  • Analyze crawl performance and web coverage to identify gaps and inefficiencies.
  • Build observability, monitoring, and reliability systems for large-scale crawl operations.
  • Debug production issues and enhance the performance, scalability, and resilience of crawling infrastructure.

Requirements

  • Experience building large-scale web crawling, search indexing, content acquisition, or internet-scale data collection systems.
  • Strong understanding of crawling architectures, URL frontier management, scheduling, and distributed crawl coordination.
  • Experience with large-scale distributed systems using frameworks like Ray, Spark, Beam, or Flink.
  • Familiarity with content extraction, HTML parsing, browser automation, rendering systems, and modern web technologies.
  • Experience operating systems that process petabyte-scale datasets.
  • Strong systems engineering skills, including reliability, observability, performance optimization, and debugging.
  • Experience designing experiments and using data to improve crawl quality, coverage, and efficiency.
  • Excellent communication skills and ability to reason about system tradeoffs and operational constraints.
  • Passion for web-scale systems and internet information collection.
  • Curiosity about web data's influence on model capabilities and willingness to iterate based on results.
  • Comfort balancing crawl quality, coverage, freshness, and operational efficiency.
  • Enjoyment working at the intersection of distributed systems, data infrastructure, and AI.

Benefits

  • Top-tier compensation (salary and equity)
  • Stock options
  • Comprehensive medical, dental, vision, and life insurance
  • Annual wellness allowance
  • Provided lunch and dinner in the office
  • 22 weeks paid parental leave
  • Unlimited paid time off (US) / 30 days paid time off (UK)
  • Visa sponsorship support
  • Regular off-sites, happy hours, and team celebrations

About Reflection ai

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.