Member of Engineering (Pre-training / Data Engineering)

Remote Remote (EMEA/East Coast) FullTime

Posted 7mo ago

Job Location

Remote (EMEA/East Coast)

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Poolside is building a world where AI drives economically valuable work and scientific progress, aiming to accelerate software development through agentic systems and frontier models. This role is a core part of the Pretraining Data team, responsible for building and scaling the Model Factory, a system for rapid training, scaling, and experimentation with foundation models. The primary mission is to architect and maintain high-performance pipelines that transform trillions of raw tokens into high-quality dataset "fuel" for models. This involves engineering ingestion, deduplication, and streaming systems for petabyte-scale data, bridging the gap between raw web crawls and GPU clusters, and directly influencing model performance through superior data modeling and distributed pipeline optimization. Collaboration with Pretraining, Posttraining, Evals, and Product teams is key to generating high-quality datasets that address missing model capabilities and downstream use cases.

Responsibilities

  • Build and maintain high-performance pipelines for trillions of tokens.
  • Deliver diverse and high-quality datasets for pre-training foundation models.
  • Collaborate with Pretraining, Posttraining, Evals, and Product teams to ensure alignment on model quality.

Requirements

  • Strong background in building production-grade, distributed data systems for machine learning.
  • Experience with orchestration tools like Slurm, Airflow, or Dagster.
  • Experience with observability and reliability tools such as CI/CD, Grafana, Prometheus.
  • Familiarity with infrastructure tools like Git, Docker, k8s, and cloud managed services.
  • Experience with batch inference (e.g., vLLM).
  • Obsession with performance, especially concerning large-scale GPU clusters and distributed pipelines.
  • Expert-level Python knowledge and ability to write clean, maintainable code.
  • Strong algorithmic foundations.
  • Proficiency with libraries like Polars, Dask, or PySpark.
  • Experience building trillion-scale SOTA pretraining datasets (nice to have).
  • Experience translating research to production at scale (nice to have).
  • Experience with OCR, web crawling, or evals (nice to have).
  • Prior experience pre-training LLMs (nice to have).

Benefits

  • Fully remote work & flexible hours
  • 37 days/year of vacation & holidays
  • Health insurance allowance for you & dependents
  • Company-provided equipment
  • Well-being, always-be-learning & home office allowances
  • Frequent team get togethers
  • Diverse & inclusive people-first culture

About poolside

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.