Distributed Training Jobs

55 open roles mentioning Distributed Training

Software Engineer - Voice Model

5d ago
x

xAI

SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. The Grok Voice Model team is building the world's best voice AI, delivering smooth, natural, low-latency spoken interactions that are expressive, multilingual, and reliable across devices and real-time scenarios. The team owns the full training pipeline, from massive data curation and premium audio processing to frontier speech-language pre-training and intensive post-training to push quality, speed, and stability to the limit. The goal is to make talking to AI feel like conversing with a charming, kind, and knowledgeable person, and exceptionally smart, execution-oriented engineers are sought to achieve this.

$150k - $450k

Palo Alto, CA onsite
KubernetesPythonFine-Tuning +5 more

Security engineer, detection and response

9d ago
W

Writer

WRITER is seeking a Staff Detection and Response Engineer to join its security team. This role is crucial for protecting the AI infrastructure that is transforming how the world works. You will be responsible for building sophisticated detection systems to identify attacks targeting our AI platform, training data, and model deployments, while also creating automated response capabilities that scale with the company's rapid growth. This position offers a unique opportunity to combine hands-on security engineering with strategic thinking to stay ahead of novel threats in cutting-edge AI/AGI systems. You will act as the operational arm of the security function, translating threat intelligence into real-time detections, coordinating incident response, and hunting for sophisticated attacks across GPU clusters and distributed training environments. This role is ideal for someone excited by the challenge of securing systems fundamentally different from traditional ones and defining the future of AI security engineering at scale.

San Francisco, CA hybrid FullTime
PythonAI AgentsDistributed Training

Security engineer, detection and response (UK)

9d ago
W

Writer

WRITER is seeking a Staff Detection and Response Engineer to join its security team and protect the company's AI infrastructure. This role involves building sophisticated detection systems for attacks targeting the AI platform, training data, and model deployments, as well as creating automated response capabilities. The position combines hands-on security engineering with strategic thinking to address novel threats in cutting-edge AI/AGI systems. The engineer will be the operational arm of the security function, translating threat intelligence into real-time detections, coordinating incident response, and hunting for sophisticated attacks across GPU clusters and distributed training environments. This is an opportunity to define AI security engineering at scale, working closely with AI Security research, Cloud Infrastructure, Software Security Engineering, and AI researchers.

London, UK hybrid FullTime
PythonAI AgentsDistributed Training

Senior Software Engineer, GPU Infrastructure (HPC)

10d ago
Cohere

Cohere

Cohere is seeking a Staff Software Engineer to join our internal infrastructure team, responsible for building and operating world-class infrastructure and tools for training, evaluating, and serving Cohere's foundational AI models. You will work closely with AI researchers to support their AI workload needs on cutting-edge systems, focusing on stability, scalability, and observability. This role involves building and operating superclusters across multiple clouds, directly accelerating the development of industry-leading AI models. Participation in a 24x7 on-call rotation is required and compensated.

Canada hybrid FullTime
CohereKubernetesPython +5 more

Internship - Machine Learning Research Engineer

12d ago
P

Perplexity AI

We are seeking motivated interns to join our Machine Learning Research Engineering team in Berlin for a 12-24 week full-time, in-person program. You will have the opportunity to significantly improve search quality by developing and optimizing large-scale deep learning models. This role involves conducting cutting-edge research in representation learning and building advanced RAG pipelines.

$12k - $24k

Berlin hybrid FullTime
RAGPyTorchDeep Learning +1 more

Research Engineer, Pretraining

17d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer to join its Pretraining team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, aiming to build safe, steerable, and trustworthy AI systems. The mission is to ensure that transformative AI systems are aligned with human interests and are beneficial for society. The team is dedicated to pushing the boundaries of AI while prioritizing safety and ethics.

London, UK onsite
AnthropicKubernetesPython +4 more

Research Scientist (Generative Modeling)

20d ago
W

World-labs

World Labs is a frontier AI research and product company focused on spatial intelligence, advancing beyond large language models. Co-founded by leading researchers, the company is pioneering world models that perceive, generate, reason, and interact with virtual and physical worlds. Their flagship product, Marble, transforms various media into navigable 3D worlds, with applications in gaming, film, architecture, robotics, and immersive experiences. Backed by significant investment, World Labs is building a world-class team at the intersection of AI research and real-world deployment.

$250k - $325k

San Francisco onsite
PythonFine-TuningPyTorch +3 more

Senior Applied Research Engineer - Video

21d ago
s

synthesia.io

Synthesia is seeking an Applied Research Engineer to join their Video team and contribute to building the next generation of production-grade foundation models for human-centric video generation. This role involves working at the intersection of large-scale generative modeling, distributed systems, and production engineering, with a focus on developing and optimizing video base models for realistic, controllable, and expressive synthetic humans. This is an applied research position with direct product impact, requiring the candidate to advance training recipes, scale distributed systems, improve evaluation frameworks, and optimize inference for real-world deployment. The work will directly influence models used by tens of thousands of businesses globally.

Europe remote FullTime
AWSDockerPython +4 more

Research Engineer, Code RL (Reinforcement Learning)

24d ago
Anthropic

Anthropic

We are seeking a Research Engineer for our Code RL team, focused on advancing AI models' capabilities in writing, editing, testing, debugging, and shipping real software. This role involves designing RL environments, coding tasks, and reward signals, as well as running training experiments on frontier models. You will diagnose model performance, improve pipeline speed and reliability, and contribute to areas like agentic coding behaviors, code correctness, and autonomous engineering. The position blends cutting-edge research with practical engineering to build high-quality, scalable AI systems.

San Francisco, CA | New York City, NY onsite
AnthropicPythonFine-Tuning +5 more

Performance Engineer, GPU

24d ago
Anthropic

Anthropic

Pioneering the next generation of AI requires breakthrough innovations in GPU performance and systems engineering. As a GPU Performance Engineer, you'll architect and implement the foundational systems that power Claude and push the frontiers of what's possible with large language models. You'll be responsible for maximizing GPU utilization and performance at unprecedented scale, developing cutting-edge optimizations that directly enable new model capabilities and dramatically improve inference efficiency. Working at the intersection of hardware and software, you'll implement state-of-the-art techniques from custom kernel development to distributed system architectures. Your work will span the entire stack—from low-level tensor core optimizations to orchestrating thousands of GPUs in perfect synchronization. Strong candidates will have a track record of delivering transformative GPU performance improvements in production ML systems and will be excited to shape the future of AI infrastructure alongside world-class researchers and engineers.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicPyTorchClaude +2 more

Research Engineer/Research Scientist, Pre-training

24d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer to join its Pre-training team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, contributing to the creation of safe, steerable, and trustworthy AI systems. The team is dedicated to ensuring that transformative AI systems are aligned with human interests and societal benefit.

Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY remote
AnthropicKubernetesPython +4 more

Staff Research Engineer, Discovery Team

24d ago
Anthropic

Anthropic

Anthropic is dedicated to building reliable, interpretable, and steerable AI systems that are safe and beneficial for society. As a Research Engineer on the Discovery Team, you will work end-to-end to identify and address key blockers on the path to scientific Artificial General Intelligence (AGI). This role involves improving models' abilities to use computers, acting as a laboratory for long-horizon tasks and a crucial component for scientific workflows. You will collaborate with a team of researchers and engineers focused on pushing the scientific frontier.

San Francisco, CA onsite
AnthropicDockerKubernetes +2 more

Member of Technical Staff, Agentic Environments

1mo ago
Cohere

Cohere

Cohere is seeking a senior engineer to join our team, focusing on the practical challenges of deploying AI systems at scale in production environments. This hands-on, engineering-driven role involves working with frontier AI models, building scalable solutions, and bridging research concepts with real-world implementations. You will contribute to both engineering and research efforts, designing and writing high-performing software for model training, developing new tools to support LLM research, and collaborating with various engineering and scientific teams. We provide access to world-class compute resources, data, and talent to enable you to do your best work.

Europe remote FullTime
CohereKubernetesPython +3 more

AI Performance Engineer

1mo ago
A

Applied Intuition

Applied Intuition is seeking a performance engineer to specialize in making large-scale machine learning workloads fast and cost-efficient within the datacenter. This role focuses on optimizing distributed training runs across multiple nodes and high-throughput batch inference for processing vast amounts of real-world autonomy logs. The primary goal is to improve throughput, cluster efficiency, and reduce cost per unit of data processed, directly impacting the company's iteration speed. You will be responsible for identifying and resolving performance bottlenecks across the entire stack, from accelerators to ML frameworks and data infrastructure, working collaboratively with various engineering teams to achieve significant improvements in training time and processing costs.

$60k - $300k

Sunnyvale onsite FullTime
KubernetesPythonGo +4 more

Research Engineer, Mid-Training

1mo ago
C

Cognition

We are an applied AI lab building end-to-end software agents, known for creating Devin, the first AI software engineer. Our team is composed of highly talented individuals with backgrounds in competitive programming and leadership roles at cutting-edge AI companies. We are tackling significant global challenges and developing AI capable of real-world reasoning. This role focuses on the critical 'mid-training' phase, bridging pre-training and post-training to refine raw model capabilities. You will be instrumental in shaping our models' fundamental abilities by owning late-stage training decisions, including data mix and quality, annealing schedules, context length extension, capability injection, and synthetic data strategies.

San Francisco onsite FullTime
PythonPyTorchDeep Learning +1 more

Principal Research Engineer, Model Training & Post-Training

1mo ago
Inflection AI

Inflection AI

Inflection AI is seeking a hands-on technical leader to own the model-improvement loop, from data and training through evaluations, post-training, release criteria, and production feedback. This role sits at the intersection of research, production engineering, and model release, with the goal of shipping measurably better models for users. The ideal candidate will have prior experience leading significant model training or post-training initiatives and can make informed tradeoffs across data, compute, architecture, and quality to achieve a clear technical roadmap.

$400k - $550k

Palo Alto, California, United States onsite
AI AgentsFine-TuningRLHF +1 more

Member of Technical Staff, AI Training Infrastructure

1mo ago
f

fireworks ai

Fireworks is seeking a Training Infrastructure Engineer to design, build, and optimize the infrastructure that powers large-scale model training operations. This role is crucial for developing high-performance AI training infrastructure, requiring collaboration with AI researchers and engineers to create robust training pipelines, optimize distributed training workloads, and ensure reliable model development. The position offers the opportunity to solve hard problems at the forefront of AI infrastructure, build what's next with bleeding-edge technology, and have a direct impact on the future of AI within a fast-growing, passionate team.

San Mateo hybrid FullTime
AWSAzureDocker +5 more

ML Research Intern

1mo ago
M

Modal

AI needs a new infrastructure layer, and we're building it at Modal. Our customers include category-defining companies like Lovable, Ramp, Cognition, DoorDash, and Suno, who rely on Modal for instant GPU access, sub-second container starts, and native storage for tasks like low-latency inference, model fine-tuning, and accessing production-ready sandboxes at scale. We are seeking PhD research interns with strong research experience in reinforcement learning, machine learning, and foundation models, including large language and multimodal models, to join our research team. This internship is ideal for candidates interested in improving existing methods and developing new techniques for large-scale model training, optimization, and inference, extending models to long-context and long-horizon tasks, and enhancing inference-time efficiency, reliability, and robustness in high-stakes real-world deployments.

New York onsite FullTime
Reinforcement LearningDistributed Training

Member of Engineering (Multimodality - Research Lead)

1mo ago
p

poolside

Poolside is building a world where AI drives economically valuable work and scientific progress, aiming to accelerate the development of Artificial General Intelligence (AGI). We focus on reshaping the developer experience with agentic systems, coding assistants, and frontier models. This role is an opportunity to build and lead a new team focused on multimodality, specifically teaching our frontier coding model to understand and process image inputs. You will shape this critical capability from its early stages, leveraging our extensive resources including a powerful model factory, thousands of GPUs, and a strong research team. The initial focus will be on image input capabilities, such as interpreting designs, generating and verifying code, and understanding diagrams, with a long-term vision of evolving towards native multimodal understanding.

Remote (EMEA/East Coast) remote FullTime
PythonTransformersDistributed Training

Software Engineer, Full Stack, Tinker

1mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking a full-stack engineer to develop and deploy the products and services that Tinker users engage with daily. This role involves working across frontend, backend, and infrastructure to build the Tinker console, developer tools, and other essential components for the platform. Tinker is a fine-tuning API that enables researchers and developers to customize frontier AI models using their own data and algorithms, managing the underlying infrastructure to provide flexibility and access to advanced capabilities.

$350k - $475k

San Francisco onsite
OpenAIMistralPython +5 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.