Member of Engineering (Pre-training / Data Engineering)
Remote • Remote (EMEA/East Coast) • FullTime
Posted 7mo ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Poolside is building a world where AI drives economically valuable work and scientific progress, aiming to accelerate software development through agentic systems and frontier models. This role is a core part of the Pretraining Data team, responsible for building and scaling the Model Factory, a system for rapid training, scaling, and experimentation with foundation models. The primary mission is to architect and maintain high-performance pipelines that transform trillions of raw tokens into high-quality dataset "fuel" for models. This involves engineering ingestion, deduplication, and streaming systems for petabyte-scale data, bridging the gap between raw web crawls and GPU clusters, and directly influencing model performance through superior data modeling and distributed pipeline optimization. Collaboration with Pretraining, Posttraining, Evals, and Product teams is key to generating high-quality datasets that address missing model capabilities and downstream use cases.
Responsibilities
- Build and maintain high-performance pipelines for trillions of tokens.
- Deliver diverse and high-quality datasets for pre-training foundation models.
- Collaborate with Pretraining, Posttraining, Evals, and Product teams to ensure alignment on model quality.
Requirements
- Strong background in building production-grade, distributed data systems for machine learning.
- Experience with orchestration tools like Slurm, Airflow, or Dagster.
- Experience with observability and reliability tools such as CI/CD, Grafana, Prometheus.
- Familiarity with infrastructure tools like Git, Docker, k8s, and cloud managed services.
- Experience with batch inference (e.g., vLLM).
- Obsession with performance, especially concerning large-scale GPU clusters and distributed pipelines.
- Expert-level Python knowledge and ability to write clean, maintainable code.
- Strong algorithmic foundations.
- Proficiency with libraries like Polars, Dask, or PySpark.
- Experience building trillion-scale SOTA pretraining datasets (nice to have).
- Experience translating research to production at scale (nice to have).
- Experience with OCR, web crawling, or evals (nice to have).
- Prior experience pre-training LLMs (nice to have).
Benefits
- Fully remote work & flexible hours
- 37 days/year of vacation & holidays
- Health insurance allowance for you & dependents
- Company-provided equipment
- Well-being, always-be-learning & home office allowances
- Frequent team get togethers
- Diverse & inclusive people-first culture