Distributed Training Jobs

39 open roles mentioning Distributed Training

Security engineer, detection and response (UK)

14h ago
W

Writer

WRITER is seeking a Staff Detection and Response Engineer to join its security team and protect the company's AI infrastructure. This role involves building sophisticated detection systems for attacks targeting the AI platform, training data, and model deployments, as well as creating automated response capabilities. The position combines hands-on security engineering with strategic thinking to address novel threats in cutting-edge AI/AGI systems. The engineer will be the operational arm of the security function, translating threat intelligence into real-time detections, coordinating incident response, and hunting for sophisticated attacks across GPU clusters and distributed training environments. This is an opportunity to define AI security engineering at scale, working closely with AI Security research, Cloud Infrastructure, Software Security Engineering, and AI researchers.

London, UK onsite FullTime
PythonAI AgentsDistributed Training +7 more

Security engineer, detection and response

14h ago
W

Writer

WRITER is seeking a Staff Detection and Response Engineer to join its security team. This role is crucial for protecting the AI infrastructure that is transforming how the world works. You will be responsible for building sophisticated detection systems to identify attacks targeting our AI platform, training data, and model deployments, while also creating automated response capabilities that scale with the company's rapid growth. This position offers a unique opportunity to combine hands-on security engineering with strategic thinking to stay ahead of novel threats in cutting-edge AI/AGI systems. You will act as the operational arm of the security function, translating threat intelligence into real-time detections, coordinating incident response, and hunting for sophisticated attacks across GPU clusters and distributed training environments. This role is ideal for someone excited by the challenge of securing systems fundamentally different from traditional ones and defining the future of AI security engineering at scale.

San Francisco, CA hybrid FullTime
PythonAI AgentsDistributed Training +9 more

Engineering Manager, GPU Infrastructure

9d ago
Cohere

Cohere

Cohere is a leading enterprise AI company building cutting-edge foundation AI models and end-to-end products. The GPU Clusters team is central to Cohere's infrastructure, responsible for building and operating the superclusters that power our frontier AI models. This role involves enabling research and development at the intersection of cutting-edge hardware, distributed systems, and AI research. As an Engineering Manager, you will lead a team of highly motivated engineers passionate about GPU infrastructure and AI, fostering a culture of technical excellence and innovation in a remote-first environment. This is a unique opportunity to shape the infrastructure powering the next generation of AI.

United States onsite FullTime
CohereKubernetesPyTorch +4 more

Software Engineer, Full Stack, Tinker

13d ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking a full-stack engineer to develop and deploy the products and services that Tinker users engage with daily. This role involves working across frontend, backend, and infrastructure to build the Tinker console, developer tools, and other essential components for the platform. Tinker is a fine-tuning API that enables researchers and developers to customize frontier AI models using their own data and algorithms, managing the underlying infrastructure to provide flexibility and access to advanced capabilities.

$350k - $475k

San Francisco onsite
OpenAIMistralPython +5 more

Site Reliability Engineer (SRE)

13d ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking a Site Reliability Engineer (SRE) to ensure the end-to-end reliability of their Tinker platform. This role involves working closely with engineers and research teams to enhance the robustness and resilience of every system layer. The SRE will be instrumental in maintaining and improving the infrastructure that supports custom AI model fine-tuning, ensuring a seamless experience for researchers and developers.

$350k - $475k

San Francisco onsite
OpenAIMistralKubernetes +3 more

Research Engineer, Knowledge Foundations

16d ago
Anthropic

Anthropic

The Knowledge Work team at Anthropic builds the training environments and evaluations that empower Claude to excel in professional workflows, including searching, analyzing, and creating content across various tools and documents. As this work expands, the underlying systems require the same rigor as the research itself. In this role, you will design and execute experiments to enhance Claude's ability to search, retrieve, and reason over information at scale. Your responsibilities will encompass environment design, data curation, RL training, evaluation, and supporting infrastructure, requiring flexibility to address progress blockers. You will collaborate closely with researchers and other RL teams to implement capabilities that directly influence Claude's performance. We believe that true ownership and impact stem from both hardening existing environments and creating new ones, ensuring the quality of the entire stack that drives superhuman epistemics.

San Francisco, CA onsite
AnthropicPythonFine-Tuning +2 more

Research Engineer, Pretraining

16d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer to join its Pretraining team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, aiming to build safe, steerable, and trustworthy AI systems. The mission is to ensure that transformative AI systems are aligned with human interests and are beneficial for society. The team is dedicated to pushing the boundaries of AI while prioritizing safety and ethics.

London, UK onsite
AnthropicKubernetesPython +4 more

Research Engineer / Research Scientist, Pre-training

16d ago
Anthropic

Anthropic

Anthropic is seeking passionate Research Scientists and Engineers to join our growing Pre-training team in Zurich. This role is at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems. You will be involved in developing the next generation of large language models, with a primary focus on multimodal capabilities, enabling LLMs to understand and interact with modalities beyond text.

Zürich, CH onsite
AnthropicKubernetesPython +2 more

Research Engineer/Research Scientist, Pre-training

16d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer to join its Pre-training team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, contributing to the creation of safe, steerable, and trustworthy AI systems. The team is dedicated to ensuring that transformative AI systems are aligned with human interests and societal benefit.

Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY remote
AnthropicKubernetesPython +4 more

Staff Research Engineer, Discovery Team

16d ago
Anthropic

Anthropic

Anthropic is dedicated to building reliable, interpretable, and steerable AI systems that are safe and beneficial for society. As a Research Engineer on the Discovery Team, you will work end-to-end to identify and address key blockers on the path to scientific Artificial General Intelligence (AGI). This role involves improving models' abilities to use computers, acting as a laboratory for long-horizon tasks and a crucial component for scientific workflows. You will collaborate with a team of researchers and engineers focused on pushing the scientific frontier.

San Francisco, CA onsite
AnthropicDockerKubernetes +2 more

Performance Engineer, GPU

16d ago
Anthropic

Anthropic

Pioneering the next generation of AI requires breakthrough innovations in GPU performance and systems engineering. As a GPU Performance Engineer, you'll architect and implement the foundational systems that power Claude and push the frontiers of what's possible with large language models. You'll be responsible for maximizing GPU utilization and performance at unprecedented scale, developing cutting-edge optimizations that directly enable new model capabilities and dramatically improve inference efficiency. Working at the intersection of hardware and software, you'll implement state-of-the-art techniques from custom kernel development to distributed system architectures. Your work will span the entire stack—from low-level tensor core optimizations to orchestrating thousands of GPUs in perfect synchronization. Strong candidates will have a track record of delivering transformative GPU performance improvements in production ML systems and will be excited to shape the future of AI infrastructure alongside world-class researchers and engineers.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicPyTorchClaude +2 more

Research Engineer, Code RL (Reinforcement Learning)

16d ago
Anthropic

Anthropic

We are seeking a Research Engineer for our Code RL team, focused on advancing AI models' capabilities in writing, editing, testing, debugging, and shipping real software. This role involves designing RL environments, coding tasks, and reward signals, as well as running training experiments on frontier models. You will diagnose model performance, improve pipeline speed and reliability, and contribute to areas like agentic coding behaviors, code correctness, and autonomous engineering. The position blends cutting-edge research with practical engineering to build high-quality, scalable AI systems.

San Francisco, CA | New York City, NY onsite
AnthropicPythonFine-Tuning +5 more

Research Scientist, Data

29d ago
pika

pika

Pika is seeking a staff or lead-level Research Engineer, Data to architect and scale data engineering systems for their advanced multimodal foundation models. This role is crucial for strengthening research teams by building, optimizing, and owning large-scale data pipelines and ML data curation. The goal is to ensure foundation models have access to high-quality, diverse datasets, enabling millions of creators. If you are passionate about data infrastructure and innovative research-engineering, this is an opportunity to make a significant impact.

Palo Alto HQ onsite FullTime
AWSAzurePython +4 more

Research Engineer, Post-Training

1mo ago
Harvey

Harvey

Harvey is transforming legal and professional services by combining agentic AI, an enterprise-grade platform, and deep domain expertise. We are seeking a Research Engineer focused on post-training to scale the process of turning expert feedback and agent traces into significantly improved models. This role involves defining and running model training experiments, interpreting results, and collaborating with internal and external partners to enhance data, environments, graders, and training methodologies. The ideal candidate is a self-manager with extensive hands-on experience training open-weight models and the engineering depth to execute and debug experiments efficiently.

$231k - $340k

San Francisco hybrid FullTime
PythonRLHFDistributed Training

ML Engineer, Inference & Optimization

1mo ago
pika

pika

We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale. You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.

Palo Alto HQ onsite FullTime
Deep LearningDistributed TrainingModel Serving

Research Engineer, Frontier Speculative Decoding

1mo ago
Together AI

Together AI

Together AI is building the Inference Platform that powers the world's most advanced generative AI models. This role will serve as a critical bridge between cutting-edge research and real-world applications, focusing on translating internal model training research into production-ready deployments for customers. The work involves a deep commitment to data-centric development, meticulous hyperparameter tuning, and rigorous checkpoint evaluation. You will transform general-purpose models into highly performant, specialized tools by fine-tuning them on customer-specific data and internal datasets, working with dedicated GPU clusters rather than training foundation models from scratch.

$190k - $270k

San Francisco, New York City remote
KubernetesPythonFine-Tuning +6 more

Research Intern, Model Shaping (Fall 2026)

1mo ago
Together AI

Together AI

As a Research Intern in the Model Shaping team, you will work on advanced post-training methods, new techniques for efficient neural network training, and robust evaluation of foundation model capabilities. The Model Shaping team at Together AI focuses on tailoring open foundation models for downstream applications, building services for machine learning developers, and developing new methods for efficient model training and evaluation. This role offers the opportunity to contribute to cutting-edge research and potentially influence open-source projects.

San Francisco hybrid
PyTorchNLPReinforcement Learning +6 more

Technical Lead Manager, Physical AI

2mo ago
Scale AI

Scale AI

Scale AI is seeking a Technical Lead Manager for its Physical AI team, focusing on the development of general AI that can reason and act in the physical world. This role bridges cutting-edge Machine Learning research with physical robot deployment, leading a team of Research Engineers while remaining a hands-on technical contributor. The primary focus is on developing and evaluating Large-Scale Foundation Models, such as VLAs and World models, to enable robots and autonomous vehicles to generalize across diverse tasks and environments. The team leverages Scale's extensive data infrastructure to help build Foundation Models for Physical AI, aiming to redefine the future of automation.

$249k - $311k

San Francisco, CA remote
Fine-TuningPyTorchReinforcement Learning +9 more

Security engineer, detection and response

2mo ago
W

Writer

WRITER is seeking a Staff Detection and Response Engineer to join its security team. This role is crucial for protecting the AI infrastructure that is transforming how the world works. You will be responsible for building sophisticated detection systems to identify attacks targeting our AI platform, training data, and model deployments, while also creating automated response capabilities that scale with the company's rapid growth. This position offers a unique opportunity to combine hands-on security engineering with strategic thinking to stay ahead of novel threats in cutting-edge AI/AGI systems. You will act as the operational arm of the security function, translating threat intelligence into real-time detections, coordinating incident response, and hunting for sophisticated attacks across GPU clusters and distributed training environments. This role is ideal for someone excited by the challenge of securing systems fundamentally different from traditional ones and defining the future of AI security engineering at scale.

US-Office Hubs hybrid FullTime
PythonAI AgentsDistributed Training

Security engineer, detection and response (UK)

2mo ago
W

Writer

WRITER is seeking a Staff Detection and Response Engineer to join its security team and protect the company's AI infrastructure. This role involves building sophisticated detection systems for attacks targeting the AI platform, training data, and model deployments, as well as creating automated response capabilities. The position combines hands-on security engineering with strategic thinking to address novel threats in cutting-edge AI/AGI systems. The engineer will be the operational arm of the security function, translating threat intelligence into real-time detections, coordinating incident response, and hunting for sophisticated attacks across GPU clusters and distributed training environments. This is an opportunity to define AI security engineering at scale, working closely with AI Security research, Cloud Infrastructure, Software Security Engineering, and AI researchers.

London, UK hybrid FullTime
PythonAI AgentsDistributed Training

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.