Member of Technical Staff - Research Engineer

€130k - €240k San Francisco (United States) FullTime

Posted 2mo ago

Job Location

San Francisco (United States)

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Research Engineer

About the job

We are seeking a Research Engineer to join our team, focusing on the critical area of large-scale model training. This role bridges the gap between cutting-edge research and production systems, tackling complex challenges in training stability, efficiency, and performance across vast GPU fleets. You will work directly with researchers, contributing code, measurements, and system changes to enable the development of foundational generative models. The ideal candidate possesses deep technical ownership, can navigate ambiguous problems, and is committed to verifying results and owning outcomes.

Responsibilities

  • Improve performance, reliability, and numerical stability of large multimodal generative model training runs.
  • Profile training steps across various components including model code, attention, kernels, data loading, and communication.
  • Implement and validate GPU-level optimizations such as fused kernels, low-precision matmuls, and quantization kernels.
  • Advance low-precision training techniques, including FP8/MXFP8/FP4 paths and quantization tradeoffs.
  • Collaborate with researchers to translate architecture changes into efficient training implementations.
  • Debug distributed training failures, including NaNs, loss spikes, numerical drift, and memory issues.
  • Build benchmarking and profiling harnesses for performance validation.
  • Address urgent bottlenecks in the training process and develop tools to prevent repeated failures.

Requirements

  • Experience with large-scale training systems, ideally in collaboration with researchers.
  • Strong PyTorch fluency, including modifying low-level training code.
  • Experience with distributed training concepts (FSDP, parallelism types, activation checkpointing, NCCL).
  • Hands-on experience improving training throughput, memory footprint, or stability.
  • Experience profiling GPU workloads using tools like Nsight Systems/Compute, torch profiler, or custom telemetry.
  • Practical GPU performance judgment and ability to verify correctness and performance.
  • Understanding of low-precision training and quantization tradeoffs (FP8, MXFP8, FP4, scaling, accumulation).
  • Good research judgment to partner with researchers and tie optimization work to model quality.
  • Comfort operating in ambiguous environments and resolving production issues.

Benefits

  • Support or co-ownership of training for a frontier foundation model.
  • Experience writing or improving GPU kernels.
  • Work on attention performance and variable sequence length training.
  • Experience with Hopper or Blackwell-class GPUs.
  • Experience with low-precision training.
  • Experience with diffusion, flow matching, DiT, and multimodal generative model training.
  • Opportunity to work on foundational generative models like Latent Diffusion and Stable Diffusion.
  • Cover reasonable travel costs for in-person collaboration.

About Black Forest Labs

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.