Member of Engineering (Synthetic Data Research)
Remote • Remote (US) • FullTime
Posted 7mo ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Research Engineer
About the job
Poolside is building a company to achieve Artificial General Intelligence, focusing on accelerating software development through agentic systems and frontier models. This role is on the data team, with the primary mission to improve the quality of datasets used for training models across the entire training cycle. The focus is on generating synthetic data at scale and determining the best strategies to leverage this data for training large models. You will collaborate with various teams to define data needs that address missing model capabilities and downstream use cases, staying current with research in synthetic data generation and LLM training. This is a hands-on role involving leading research initiatives through experiments and deploying engineering solutions into production, with access to performant distributed data pipelines and large GPU clusters.
Responsibilities
- Stay updated on the latest research in LLMs and synthetic data generation, and be familiar with relevant open-source datasets and models.
- Design and implement complex pipelines for generating large, diverse datasets efficiently.
- Collaborate with cross-functional teams to ensure experiments and data generation are resource-efficient and improve model quality.
- Continuously measure and refine dataset quality, validating data strategies through quantitative ablation experiments.
Requirements
- Strong machine learning and engineering background.
- Experience with Large Language Models (LLMs), including understanding of LLM learning, data ablations, scaling laws, post-training techniques, and training reasoning/agentic models.
- Experience implementing cost-efficient, complex pipelines for large-scale synthetic dataset generation, optimizing for quality, correctness, and diversity.
- Experience with evals for tracking model capabilities (general knowledge, reasoning, math, coding, long-context).
- Experience building trillion-scale pretraining datasets, including familiarity with data curation, deduplication, mixing, tokenization, curriculum, and data repetition impact.
- Excellent programming skills in Python.
- Strong prompt engineering skills.
- Experience with large-scale GPU clusters and distributed data pipelines.
- Strong obsession with data quality.
- Research experience, ideally with publications in applied deep learning, LLMs, or source code generation.
- Ability to discuss the latest papers in detail and hold informed opinions.
Benefits
- Fully remote work & flexible hours
- 37 days/year of vacation & holidays
- Health insurance allowance