Research Scientist, Data

Palo Alto HQ FullTime

Posted 29d ago

Job Location

Palo Alto HQ

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Research Engineer

About the job

Pika is seeking a staff or lead-level Research Engineer, Data to architect and scale data engineering systems for their advanced multimodal foundation models. This role is crucial for strengthening research teams by building, optimizing, and owning large-scale data pipelines and ML data curation. The goal is to ensure foundation models have access to high-quality, diverse datasets, enabling millions of creators. If you are passionate about data infrastructure and innovative research-engineering, this is an opportunity to make a significant impact.

Responsibilities

  • Own large-scale data pipeline architecture and implementation for model training and research workflows (text, image, audio, video).
  • Partner with research and engineering teams to curate, clean, and manage diverse datasets for multimodal model pre-training and mid-training.
  • Develop strategies and tools for scalable data ingestion, labeling, filtering, augmentation, and storage.
  • Ensure data quality, reliability, and compliance, including privacy and ethical considerations.
  • Optimize data processing, transformation, and delivery for large-scale distributed training pipelines.
  • Prototype and productionize new methods for dataset creation, management, and continuous improvement.
  • Contribute to integrating research-driven data advancements into production systems.
  • Stay informed on emerging data engineering and ML data management developments and apply best practices.

Requirements

  • 5+ years of experience building and scaling data pipelines for machine learning applications at staff or lead engineer level, ideally in research or model training environments.
  • Strong background in data engineering and ML data curation for LLMs, VLMs, or other large-scale multimodal models.
  • Expertise in distributed data systems (e.g., Spark, Hadoop, Ray) and efficient large dataset processing/ETL workflows.
  • Proven ability to build robust, scalable, and production-grade data infrastructure for ML pipelines.
  • Experience developing tools for data labeling, filtering, deduplication, quality assurance, and dataset management.
  • Strong programming skills (Python, SQL, PySpark) and familiarity with cloud data platforms (AWS, GCP, Azure).
  • Knowledge of privacy, compliance, ethics, and best practices in data collection and management.
  • Excellent cross-functional collaboration, problem-solving, and communication skills.
  • Passion for enabling cutting-edge generative AI and creative technology through data excellence.

Benefits

  • Competitive salary
  • Substantial equity in a high-growth startup
  • Full health benefits
  • 401k matching
  • Collaborative, mission-driven team environment
  • Major growth opportunities

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.