Distributed Training Jobs

55 open roles mentioning Distributed Training

Site Reliability Engineer (SRE)

1mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking a Site Reliability Engineer (SRE) to ensure the end-to-end reliability of their Tinker platform. This role involves working closely with engineers and research teams to enhance the robustness and resilience of every system layer. The SRE will be instrumental in maintaining and improving the infrastructure that supports custom AI model fine-tuning, ensuring a seamless experience for researchers and developers.

$350k - $475k

San Francisco onsite
OpenAIMistralKubernetes +3 more

Staff Research Engineer - Multimodal Generative Modelling

1mo ago
s

synthesia.io

Synthesia is seeking a Staff Research Engineer to join their Voice team and contribute to the company's long-term vision of building the best human-interactive models. These advanced systems will go beyond simple conversation to perceive, respond to, and react to user actions and emotions in real-time. The role involves defining and driving this broader vision across teams, proposing ambitious research directions, and taking ownership of critical component design and implementation. You will collaborate closely with the voice lead, video teams, and other senior members to develop models that combine text, audio, and video for a seamless, interactive experience, moving beyond traditional turn-taking in speech models.

Europe remote FullTime
Fine-TuningPyTorchDeep Learning +1 more

Principal Product Manager - Hardware

2mo ago
L

Lambda

Lambda is seeking a Principal Product Manager to own the hardware product lifecycle for their AI cloud infrastructure. This role is critical in deciding which GPU platforms, node configurations, and cluster options Lambda offers to its customers. You will work at the intersection of hardware and software, collaborating with external partners like NVIDIA and ODMs, as well as internal teams including data center, supply chain, and infrastructure engineering. Your responsibilities will include translating customer demand and benchmark data into strategic fleet investment recommendations, defining product roadmaps, and ensuring the successful launch and iteration of hardware products. The ideal candidate possesses strong product management skills, a deep understanding of technical infrastructure and hardware, and the ability to drive decisions through insight, influence, and execution.

Bellevue Office hybrid FullTime
Distributed Training

Research Scientist, Data

2mo ago
pika

pika

Pika is seeking a staff or lead-level Research Engineer, Data to architect and scale data engineering systems for their advanced multimodal foundation models. This role is crucial for strengthening research teams by building, optimizing, and owning large-scale data pipelines and ML data curation. The goal is to ensure foundation models have access to high-quality, diverse datasets, enabling millions of creators. If you are passionate about data infrastructure and innovative research-engineering, this is an opportunity to make a significant impact.

Palo Alto HQ onsite FullTime
AWSAzurePython +4 more

Member of Technical Staff - Research Engineer

2mo ago
B

Black Forest Labs

We are seeking a Research Engineer to join our team, focusing on the critical area of large-scale model training. This role bridges the gap between cutting-edge research and production systems, tackling complex challenges in training stability, efficiency, and performance across vast GPU fleets. You will work directly with researchers, contributing code, measurements, and system changes to enable the development of foundational generative models. The ideal candidate possesses deep technical ownership, can navigate ambiguous problems, and is committed to verifying results and owning outcomes.

€130k - €240k

San Francisco (United States) onsite FullTime
PyTorchDistributed Training

Research Engineer, Post-Training

2mo ago
Harvey

Harvey

Harvey is transforming legal and professional services by combining agentic AI, an enterprise-grade platform, and deep domain expertise. We are seeking a Research Engineer focused on post-training to scale the process of turning expert feedback and agent traces into significantly improved models. This role involves defining and running model training experiments, interpreting results, and collaborating with internal and external partners to enhance data, environments, graders, and training methodologies. The ideal candidate is a self-manager with extensive hands-on experience training open-weight models and the engineering depth to execute and debug experiments efficiently.

$231k - $340k

San Francisco hybrid FullTime
PythonRLHFDistributed Training

ML Engineer, Inference & Optimization

2mo ago
pika

pika

We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale. You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.

Palo Alto HQ onsite FullTime
Deep LearningDistributed TrainingModel Serving

Research Intern, Model Shaping (Fall 2026)

3mo ago
Together AI

Together AI

As a Research Intern in the Model Shaping team, you will work on advanced post-training methods, new techniques for efficient neural network training, and robust evaluation of foundation model capabilities. The Model Shaping team at Together AI focuses on tailoring open foundation models for downstream applications, building services for machine learning developers, and developing new methods for efficient model training and evaluation. This role offers the opportunity to contribute to cutting-edge research and potentially influence open-source projects.

San Francisco hybrid
PyTorchNLPReinforcement Learning +6 more

Research Engineer, Frontier Speculative Decoding

3mo ago
Together AI

Together AI

Together AI is building the Inference Platform that powers the world's most advanced generative AI models. This role will serve as a critical bridge between cutting-edge research and real-world applications, focusing on translating internal model training research into production-ready deployments for customers. The work involves a deep commitment to data-centric development, meticulous hyperparameter tuning, and rigorous checkpoint evaluation. You will transform general-purpose models into highly performant, specialized tools by fine-tuning them on customer-specific data and internal datasets, working with dedicated GPU clusters rather than training foundation models from scratch.

$190k - $270k

San Francisco, New York City remote
KubernetesPythonFine-Tuning +6 more

Technical Lead Manager, Physical AI

3mo ago
Scale AI

Scale AI

Scale AI is seeking a Technical Lead Manager for its Physical AI team, focusing on the development of general AI that can reason and act in the physical world. This role bridges cutting-edge Machine Learning research with physical robot deployment, leading a team of Research Engineers while remaining a hands-on technical contributor. The primary focus is on developing and evaluating Large-Scale Foundation Models, such as VLAs and World models, to enable robots and autonomous vehicles to generalize across diverse tasks and environments. The team leverages Scale's extensive data infrastructure to help build Foundation Models for Physical AI, aiming to redefine the future of automation.

$249k - $311k

San Francisco, CA remote
Fine-TuningPyTorchReinforcement Learning +9 more

Principal Applied Research Engineer

3mo ago
s

synthesia.io

Synthesia is seeking a Principal Research Engineer to lead the technical direction for offline video generation. This role involves owning the end-to-end process, from pre-training to post-training, and resolving the complexities that arise at scale. You will partner with research leadership to define long-term strategy, tackle challenging technical problems, and accelerate the delivery of research into production. The ideal candidate has hands-on experience training large generative models from scratch and is driven by a passion for pushing the boundaries of video generation and ensuring that work reaches users.

Europe remote FullTime
Fine-TuningRLHFDistributed Training

ML Researcher, Foundational Models

3mo ago
s

sarvam

Sarvam is building the bedrock of Sovereign AI for India, developing a full-stack AI platform focused on making AI genuinely work for India. This role is for a researcher who will tackle open-ended questions about the architecture, optimization, data composition, and training dynamics of our next generation of foundational models. You will have direct access to large compute resources and a tight feedback loop with engineers, driving research from initial hunches to production-ready decisions. This is a hands-on role requiring independent research design, execution at scale, and the ability to translate findings into concrete proposals for production training runs.

Bengaluru onsite FullTime
Fine-TuningPyTorchTransformers +2 more

Research, Pre-Training Data

4mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking pre-training researchers to join their mission of advancing collaborative general intelligence. This role is central to developing the next generation of AI models by blending research with large-scale data engineering. You will be responsible for assembling pre-training datasets and data systems, designing and implementing methods for sourcing, curating, and analyzing data for quality and performance. The position involves working with automated pipelines and human-in-the-loop processes, contributing both scientific insights and production-grade code. It's an ideal opportunity for individuals passionate about the intersection of data, machine learning, and systems, and who are eager to shape the future of AI.

$350k - $475k

San Francisco onsite
OpenAIMistralPython +5 more

Research Engineer, Infrastructure, Numerics

4mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking an infrastructure research engineer to design and build core systems for efficient large-scale model training, with a specific focus on numerics. This role involves enhancing the numerical foundations of their distributed training stack, optimizing precision formats, kernel optimizations, and communication frameworks to ensure stable, scalable, and fast training of trillion-parameter models. The ideal candidate will bridge research and systems engineering, possessing a strong understanding of both optimization mathematics and distributed compute realities.

$350k - $475k

San Francisco onsite
OpenAIMistralPyTorch +3 more

Member of Engineering (Post-training)

4mo ago
p

poolside

Poolside is building a company to create Artificial General Intelligence, aiming to accelerate software development through agentic systems, coding assistants, and frontier models. This role is part of the Applied Research team, focused on transforming pre-trained Large Language Models (LLMs) into well-aligned and highly capable AI systems specifically for coding and software development. You will be involved in building data pipelines and environments for agentic use cases, researching and implementing post-training algorithms, and designing experiments to test hypotheses, with access to significant GPU resources.

Remote (EMEA/East Coast) remote FullTime
PythonFine-TuningPyTorch +4 more

Member of Technical Staff - VLM

4mo ago
B

Black Forest Labs

Black Forest Labs is seeking a Staff / Senior Individual Contributor to pioneer the integration of vision-language models (VLMs) directly into their advanced generative AI systems, like FLUX. This role focuses on developing novel VLM approaches and architectures, rather than just applying existing ones. You will explore how vision and language representations inform each other, how multimodal understanding enhances generation quality, and how to deploy these capabilities at scale. The ideal candidate will have a proven track record of pretraining or significantly advancing a VLM that has been deployed or publicly released.

€130k - €340k

Freiburg (Germany) onsite FullTime
Fine-TuningDistributed Training

Member of Technical Staff - Pretraining

4mo ago
B

Black Forest Labs

We are seeking a Member of Technical Staff focused on pretraining to join our research team. This role is central to developing the next generation of multimodal foundation models for image, video, and audio. You will have a significant impact on shaping training objectives, architectures, data strategies, and systems, with your research directly influencing products used by millions. This is a Staff/Senior individual contributor position for someone with proven experience leading frontier pretraining efforts.

€130k - €340k

Freiburg (Germany) onsite FullTime
PythonPyTorchDistributed Training

Member of Technical Staff (AI Infrastructure Engineer)

5mo ago
P

Perplexity AI

We are seeking an AI Infrastructure Engineer to join our expanding team. In this role, you will collaborate closely with our Inference and Research teams to construct, deploy, and enhance our extensive AI training and inference clusters. Your work will involve managing Kubernetes and Slurm environments, optimizing distributed training for large language models, and developing robust orchestration systems. This position offers the opportunity to significantly impact the performance and scalability of our AI infrastructure.

London hybrid FullTime
AWSKubernetesPython +5 more

Member of Technical Staff (AI Infrastructure Engineer)

5mo ago
P

Perplexity AI

We are seeking an AI Infrastructure Engineer to join our expanding team. In this role, you will collaborate closely with our Inference and Research teams to construct, deploy, and enhance our large-scale AI training and inference clusters. Our work involves Kubernetes, Slurm, Python, C++, PyTorch, and primarily operates on AWS.

San Francisco onsite FullTime
AWSKubernetesPython +5 more

Software Engineer, Workload Enablement

5mo ago
OpenAI

OpenAI

OpenAI is seeking a Software Engineer to join the Scaling team, which builds the architectural and engineering backbone for the company's infrastructure. This role focuses on enabling production workloads and end-to-end testing on new platforms. You will be responsible for creating test harnesses, developing platform stress benchmarks, and porting existing AI model training and inference workloads to new systems and hardware. A key aspect of this position involves analyzing performance, identifying bottlenecks, and characterizing the behavior of new compute, communication, storage, and control plane systems, including their failure modes.

San Francisco hybrid FullTime
OpenAIKubernetesPython +3 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.