Software Engineer, Pretraining

San Francisco FullTime

Posted 20d ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

We are seeking Software Engineers to develop the data systems essential for training our frontier coding models. This role involves working on large-scale crawling, data platform, and pipeline infrastructure to transform raw data into training datasets, ensuring rapid and reliable iteration for researchers. You will contribute to building models, high-performance pipelines, and conducting experiments to validate data quality and drive model performance improvements.

Responsibilities

  • Build and own high-throughput, telemetered data pipelines with end-to-end traceability for frontier-scale data processing.
  • Train and deploy models for data classification, ranking, filtering, and cleaning at high throughput.
  • Design and execute scaling-ladder experiments on data mixtures, repeatability, and quality depth.
  • Partner with Data Acquisition and training teams to address data source issues and optimize model performance.
  • Treat data quality as a systems and research problem, writing performance-critical code and designing experiments.
  • Build the platform for transforming raw web, code, multimodal, and acquired data into training-ready datasets.
  • Own pipelines, orchestration, and tooling for fast, reliable, observable, and reproducible pretraining data iteration.
  • Create clear signals for data quality, lineage, freshness, and pipeline health.
  • Partner with cross-functional teams to translate new data ideas into measurable model improvements.
  • Build and scale web crawling systems for discovering, fetching, and parsing high-quality documents.
  • Improve URL seeding, scoring, and host scheduling for optimal crawl capacity.
  • Enhance crawl success and parsing quality, overcoming anti-bot measures and improving content extraction.
  • Debug and harden complex crawl infrastructure for availability, recovery, and ingestion lag.
  • Automate the delivery of crawl datasets into the data pipeline.
  • Work independently and alongside AI agents to ensure new coverage translates to improved training data.

Requirements

  • Strong infrastructure or data platform background.
  • Experience in crawling or search infrastructure is a plus.
  • Demonstrated ability to move fast and take on significant ownership early in a career, or deep domain experience.
  • Ability to architect and ship end-to-end solutions with high ownership.
  • Proficiency in debugging complex systems independently.
  • Strong intuitions about large-scale distributed systems.
  • Excitement to learn how pre-training data shapes model quality and to build systems feeding frontier training runs.

About cursor

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.