Member of Technical Staff - Engineering Lead, Data Ingestion

San Francisco, CA FullTime

Posted 10d ago

Job Location

San Francisco, CA

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Reflection's Data team is responsible for building the training corpora that our frontier models learn from. The ingestion layer is the critical machinery that transforms raw data from the open web, licensed sources, and other large-scale origins into well-structured, versioned, and auditable datasets for pre-training. As the Data Ingestion Lead, you will provide front-line leadership for the team responsible for this layer, overseeing web crawl, data ingestion pipelines, and data lakes. You will build, mentor, and grow a team of data ingestion engineers, guide technical and architectural decisions across crawling, extraction, and corpus storage/delivery, and collaborate closely with research, data quality, and data partnerships teams. You will also remain hands-on to contribute technically and maintain a deep understanding of the team's work, as subtle decisions at the ingestion layer significantly impact model performance, safety, and failure points.

Responsibilities

  • Build, mentor, and grow a high-performing team of data ingestion engineers, supporting their professional development.
  • Lead the full ingestion stack, including web crawling, data extraction and normalization pipelines, and data lakes for training corpora.
  • Contribute technically to the team's stack to maintain a deep understanding and make targeted contributions.
  • Manage day-to-day execution, prioritizing work and running data acquisition campaigns in a dynamic environment.
  • Guide technical and architectural decisions, focusing on scalability, reliability, auditability, and cost for crawler architecture, distributed processing, orchestration, storage formats, deduplication, and dataset versioning.
  • Collaborate with pre-training and data quality teams to link ingestion decisions to measurable model impact and evaluate crawling and extraction strategies.
  • Work with data partnerships, digitization operations, and external vendors to onboard new data sources while adhering to legal and licensing constraints.
  • Elevate technical judgment, prioritization, communication, and execution within a fast-paced environment.

Requirements

  • Experience building, mentoring, and growing data or infrastructure engineering teams, with a willingness to scale leadership.
  • Deep experience building web-scale data acquisition or ingestion systems with ownership of production-grade pipelines at multi-TB to PB scale.
  • Strong coding ability and technical credibility.
  • Deep expertise in web crawling & acquisition, large-scale extraction & ingestion pipelines, or data lakes / corpus storage & delivery, with working knowledge across the others.
  • Fluency with modern large-scale data tools, including distributed compute (Ray, Beam, Spark), orchestration (Airflow, Prefect), formats (Parquet, JSONL, WARC), and object-store/data-lake architectures.
  • Familiarity with LLM training and evaluation, and an understanding of data's impact on model performance.
  • Ability to guide strategy, drive execution across data campaigns, and partner effectively with research, vendors, and legal teams.

Benefits

  • Top-tier compensation (salary and equity)
  • Stock options
  • Comprehensive medical, dental, vision, and life insurance
  • Annual wellness allowance
  • Provided lunch and dinner in the office
  • 22 weeks paid parental leave
  • Unlimited paid time off (US) / 30 days paid time off (UK)
  • Visa sponsorship support
  • Regular off-sites, happy hours, and team celebrations

About Reflection ai

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.