Member of Engineering (Pre-training / Data Research)

Remote Remote (EMEA/East Coast) FullTime

Posted 3mo ago

Job Location

Remote (EMEA/East Coast)

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

AI Research Engineer

About the job

Poolside is building a company to create Artificial General Intelligence, aiming to accelerate software development through agentic systems, coding assistants, and frontier models. This role focuses on improving the quality of pretraining datasets for our models. You will leverage your experience and conduct training experiments to enhance dataset quality, including synthetic data generation and data mix optimization. Collaboration with Pretraining, Posttraining, Evals, and Product teams is essential to define data needs that address missing model capabilities and downstream use cases. Staying current with research in dataset design and pretraining is crucial, as you will lead original research initiatives and deploy technical engineering solutions into production, utilizing a performant distributed data pipeline and a large GPU cluster.

Responsibilities

  • Follow the latest research related to LLMs and data quality, and be familiar with relevant open-source datasets and models.
  • Design and implement complex pipelines for generating large, diverse datasets while optimizing resource usage.
  • Collaborate with Pretraining, Posttraining, Evals, and Product teams to ensure rapid feedback loops on model quality.
  • Suggest, conduct, and analyze data ablations or training experiments to improve generated dataset quality using quantitative insights.

Requirements

  • Strong machine learning and engineering background.
  • Experience with Large Language Models (LLMs), including understanding transformer architectures, how LLMs learn, data ablations, scaling laws, mid-training and post-training techniques, and training reasoning/agentic models.
  • Experience with evals for tracking model capabilities (general knowledge, reasoning, math, coding, long-context, etc.).
  • Experience building trillion-scale pretraining datasets, including data curation, deduplication, data mixing, tokenization, curriculum, and the impact of data repetition.
  • Excellent programming skills in Python.
  • Strong prompt engineering skills.
  • Experience with large-scale GPU clusters and distributed data pipelines.
  • Strong obsession with data quality.
  • Research experience, ideally authoring scientific papers in applied deep learning, LLMs, or source code generation, and ability to discuss current research papers in detail.

Benefits

  • Fully remote work & flexible hours
  • 37 days/year of vacation & holidays
  • Health insurance allowance for you & dependents
  • Company-provided equipment
  • Well-being, always-be-learning & home office allowances
  • Frequent team get togethers
  • Diverse & inclusive people-first culture

About poolside

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.