Distributed Training Jobs

39 open roles mentioning Distributed Training

Research, Pre-Training Data

2mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking pre-training researchers to join their mission of advancing collaborative general intelligence. This role is central to developing the next generation of AI models by blending research with large-scale data engineering. You will be responsible for assembling pre-training datasets and data systems, designing and implementing methods for sourcing, curating, and analyzing data for quality and performance. The position involves working with automated pipelines and human-in-the-loop processes, contributing both scientific insights and production-grade code. It's an ideal opportunity for individuals passionate about the intersection of data, machine learning, and systems, and who are eager to shape the future of AI.

$350k - $475k

San Francisco onsite
OpenAIMistralPython +5 more

Research Engineer, Infrastructure, Numerics

2mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking an infrastructure research engineer to design and build core systems for efficient large-scale model training, with a specific focus on numerics. This role involves enhancing the numerical foundations of their distributed training stack, optimizing precision formats, kernel optimizations, and communication frameworks to ensure stable, scalable, and fast training of trillion-parameter models. The ideal candidate will bridge research and systems engineering, possessing a strong understanding of both optimization mathematics and distributed compute realities.

$350k - $475k

San Francisco onsite
OpenAIMistralPyTorch +3 more

Senior Member of Technical Staff, Safety and Security for Agents

4mo ago
Cohere

Cohere

Cohere is seeking a Senior Member of Technical Staff to join the Safety and Security for Agents team. In this role, you will significantly contribute to the development of safer, fairer, more trustworthy, and more secure Large Language Models (LLMs). Your work will focus on data generation, post-training algorithms, and evaluation methods to ensure the safety of next-generation models that interact with external resources and take actions. You will collaborate closely with machine learning teams, data annotation teams, and product and policy teams, requiring a blend of machine learning expertise, ethical AI principles, experimental design, and data management skills. This position offers significant autonomy and decision-making power within a small team, with the opportunity to shape the future of LLMs for societal benefit.

London hybrid FullTime
CoherePythonPyTorch +3 more

Research Engineer

5mo ago
Cohere

Cohere

Cohere Labs is seeking Research Engineers to join their dedicated research arm, focused on pushing machine learning forward through open, collaborative research and hands-on experimentation. This is a highly practical role where you will work closely with scientists and engineers to implement new methods, run large-scale experiments, and help shape the infrastructure supporting our research programs. You will be responsible for building experiments, debugging models, scaling training pipelines, and turning research ideas into working systems. We value curiosity, strong fundamentals, and a willingness to learn quickly in a fast-moving research environment, with a focus on practical impact.

Toronto remote FullTime
CohereFine-TuningPyTorch +4 more

Research Engineer, Machine Learning

6mo ago
Mistral AI

Mistral AI

Mistral AI is democratizing AI through high-performance, optimized, open-source models and solutions. We are a dynamic, collaborative team passionate about AI's potential to transform society, with a diverse workforce driving innovation. As a Research Engineer – ML track, you will build and optimize large-scale learning systems powering our open-weight models. You will work hand-in-hand with Research Scientists, either enhancing the shared training framework and data pipelines or embedding within a research squad to turn fresh ideas into scalable code.

Palo Alto remote Full-time
MistralPythonPyTorch +9 more

Software Engineer, GPU Infrastructure (HPC)

6mo ago
Cohere

Cohere

Cohere is seeking a Staff Software Engineer to join our internal infrastructure team, responsible for building and operating world-class infrastructure and tools for training, evaluating, and serving Cohere's foundational AI models. You will work closely with AI researchers to support their AI workload needs on cutting-edge systems, focusing on stability, scalability, and observability. This role involves building and operating superclusters across multiple clouds, directly accelerating the development of industry-leading AI models. Participation in a 24x7 on-call rotation is required and compensated.

Canada hybrid FullTime
CohereKubernetesPython +5 more

Member of Technical Staff, Senior/Staff MLE

6mo ago
Cohere

Cohere

Cohere is seeking a Member of Technical Staff, Applied ML to work directly with enterprise customers on challenging problems that push Large Language Models (LLMs) to their limits. In this role, you will rapidly understand customer domains, design custom LLM solutions, and deliver production-ready models that solve high-value, real-world problems. You will train and customize frontier models, leveraging Cohere's full stack including CPT, post-training, retrieval and agent integrations, model evaluations, and state-of-the-art modeling techniques. Your work will directly influence the capabilities of Cohere's foundation models, as techniques, datasets, evaluations, and insights developed for customers will shape the next generation of models. This role offers an opportunity to operate with early-startup level ownership within a frontier-model company, combining the breadth of an early-stage CTO with the infrastructure and scale of a deep-learning lab, and wearing multiple hats to define Applied ML at Cohere.

San Francisco remote FullTime
CoherePythonDistributed Training

Member of Technical Staff, MLE

6mo ago
Cohere

Cohere

Cohere is seeking a Member of Technical Staff, Applied ML to work directly with enterprise customers on challenging problems that push the limits of Large Language Models (LLMs). In this role, you will gain a deep understanding of customer domains, design custom LLM solutions, and deliver production-ready models to solve high-value business problems. You will not just use APIs, but will train and customize frontier models using Cohere's full stack, including CPT, post-training, retrieval and agent integrations, model evaluations, and state-of-the-art modeling techniques. Your work will directly influence the capabilities of Cohere's foundation models, shaping their next generation. This role offers a unique opportunity to combine the breadth of an early-stage CTO with the infrastructure and scale of a deep-learning lab, requiring you to wear multiple hats, set a high technical bar, and define the future of Applied ML at Cohere.

San Francisco remote FullTime
CoherePythonDistributed Training

Member of Technical Staff, Data Analysis and Evaluation

7mo ago
Cohere

Cohere

Cohere is seeking a Member of Technical Staff in Data Analysis and Evaluation to ensure the quality, reliability, and performance of our large language models (LLMs). This role involves designing and conducting data collection tasks, assessing dataset quality, and analyzing model robustness and generalisability. You will collaborate with researchers, engineers, and data annotators to drive data-driven decisions and enhance AI system effectiveness. The position requires expertise in statistics, experimental design, and machine learning to ensure high-quality data and reliable model performance across diverse scenarios, contributing to Cohere's mission of advancing AI.

London remote FullTime
CoherePythonPyTorch +3 more

Senior ML Systems Engineer, Frameworks & Tooling

8mo ago
Cohere

Cohere

Cohere is seeking a Senior ML Systems Engineer to join their team and build, maintain, and evolve the training framework that powers their frontier-scale language models. This role is ideal for someone passionate about large-scale training, distributed systems, and HPC infrastructure, offering the opportunity to design and maintain core components for fast, reliable, and scalable model training. You will also build tooling to connect research ideas to thousands of GPUs, working across the full stack of ML systems with significant autonomy and impact.

London remote FullTime
CohereDockerKubernetes +3 more

Forward Deployed Engineer, Lead - LLM Post-training

9mo ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible. We are seeking an exceptional technical leader to build and scale our post-training and evaluation capabilities within the Applied AI team. This role involves taking our open-weight models and adapting them for specific customer domains, tasks, and constraints. You will own the end-to-end technical strategy for model customization, from synthetic data generation and reward modeling through training and production deployment, working directly with customers and research teams.

New York, NY onsite FullTime
Reinforcement LearningDistributed Training

Member of Technical Staff - Pre-Training

9mo ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible for everyone. We build open models that empower individuals to control their intelligence and shape the future of AI. As a Member of Technical Staff - Pre-Training, you will be instrumental in researching and building solutions across algorithms, scaling laws, data processing, optimizers, and model architecture. This role involves designing and executing scientific experiments to deepen our understanding of scaling large language models and data efficiency, while also implementing state-of-the-art deep learning methods. You will have the opportunity to lead small research projects independently and contribute to larger initiatives, optimizing training infrastructure for efficient scaling and working across the entire stack from low-level optimizations to high-level model design.

San Francisco, CA onsite FullTime
PythonPyTorchDeep Learning +1 more

Member of Technical Staff, MLE (UK/EU)

10mo ago
Cohere

Cohere

Cohere is seeking a Member of Technical Staff for their Applied ML team. In this role, you will collaborate directly with customers to understand their challenges and implement solutions leveraging Large Language Models. You will apply your problem-solving skills, creativity, and technical expertise to bridge the gap in enterprise AI adoption, delivering impactful products and disrupting key industries. This is an opportunity to join at a pivotal moment, shape the company's offerings, and contribute to cutting-edge AI development.

London hybrid FullTime
CoherePythonTensorFlow +2 more

Machine Learning Infrastructure Engineer, Model Inference

11mo ago
Abridge

Abridge

Abridge is seeking an ML Infrastructure Engineer, Model Inference to build and optimize the core inference infrastructure powering their machine learning models. This role is crucial for enhancing the scalability, efficiency, and performance of Abridge's AI-driven healthcare solutions. The engineer will collaborate with Infrastructure and Research teams to build, deploy, optimize, and orchestrate AI models, working on a platform that transforms patient-clinician conversations into structured clinical notes in real-time.

SF Office hybrid FullTime
KubernetesPyTorchTensorFlow +3 more

Member of Technical Staff, Integration/RL Team (Research Engineer)

11mo ago
Cohere

Cohere

Cohere is a leading enterprise AI company focused on building cutting-edge foundation AI models and end-to-end products for real-world business problems. The integration team specifically focuses on developing and scaling machine learning algorithms and infrastructure for LLM post-training, with an emphasis on large-scale, distributed Reinforcement Learning (RL) methods. This role is crucial for enhancing the post-training codebase by implementing new research tools, optimizing algorithms, and scaling distributed RL capabilities. We are looking for passionate individuals who are meticulous in their approach to engineering and science, contributing to both production code and research efforts.

Paris remote FullTime
CohereKubernetesPython +3 more

Member of Technical Staff, Post-Training

1y ago
Cohere

Cohere

Cohere is seeking a Member of Technical Staff to focus on post-training of AI models. This role is crucial for advancing the state of the art in model post-training and shipping cutting-edge models to production, bridging the gap between research and practical application. You will have access to significant compute resources and a talented team to contribute to increasing model capabilities and driving customer value. The position offers a unique opportunity to contribute to both production code and research efforts, depending on your interests and organizational needs.

London hybrid FullTime
CohereKubernetesPython +3 more

Member of Technical Staff, Search

1y ago
Cohere

Cohere

Cohere is seeking talented individuals to join our Search team as a Member of Technical Staff. This role focuses on developing state-of-the-art models for information retrieval, including training embedding and reranker models. You will have the opportunity to revolutionize search experiences by building intelligent, efficient, and precise search systems. Your work will advance semantic search techniques, improve accuracy and efficiency, and involve integrating novel technologies into our search infrastructure. We are looking for someone passionate about search with a strong background in information retrieval and experience working with diverse technologies and cross-functional teams.

United States remote FullTime
CoherePythonGo +5 more

Member of Technical Staff, MLE (Korea)

1y ago
Cohere

Cohere

Cohere is seeking a Member of Technical Staff for their Applied ML team. In this role, you will collaborate directly with customers to understand their challenges and implement solutions leveraging Large Language Models. You will apply your problem-solving skills, creativity, and technical expertise to bridge the gap in Enterprise AI adoption, delivering impactful products and disrupting key industries. This is an opportunity to join at a pivotal moment, shape the company's offerings, and contribute across multiple facets of the business.

Korea remote FullTime
CoherePythonTensorFlow +2 more

Research Engineer, Machine Learning - Paris/London/Zurich/Warsaw

1y ago
Mistral AI

Mistral AI

Mistral AI is a pioneering company focused on democratizing AI through high-performance, optimized, open-source models and solutions. We aim to simplify tasks, save time, and enhance learning and creativity by integrating AI seamlessly into daily working life. Our comprehensive AI platform serves both enterprise and personal needs, featuring offerings like Le Chat, La Plateforme, Mistral Code, and Mistral Compute. We are a dynamic, collaborative, and diverse team passionate about AI's potential to transform society, driven by innovation and a low-ego, team-spirited culture.

Paris hybrid Full-time
MistralPythonPyTorch +11 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.