Distributed Training Jobs
39 open roles mentioning Distributed Training
Security engineer, detection and response (UK)
Writer
WRITER is seeking a Staff Detection and Response Engineer to join its security team and protect the company's AI infrastructure. This role involves building sophisticated detection systems for attacks targeting the AI platform, training data, and model deployments, as well as creating automated response capabilities. The position combines hands-on security engineering with strategic thinking to address novel threats in cutting-edge AI/AGI systems. The engineer will be the operational arm of the security function, translating threat intelligence into real-time detections, coordinating incident response, and hunting for sophisticated attacks across GPU clusters and distributed training environments. This is an opportunity to define AI security engineering at scale, working closely with AI Security research, Cloud Infrastructure, Software Security Engineering, and AI researchers.
Security engineer, detection and response
Writer
WRITER is seeking a Staff Detection and Response Engineer to join its security team. This role is crucial for protecting the AI infrastructure that is transforming how the world works. You will be responsible for building sophisticated detection systems to identify attacks targeting our AI platform, training data, and model deployments, while also creating automated response capabilities that scale with the company's rapid growth. This position offers a unique opportunity to combine hands-on security engineering with strategic thinking to stay ahead of novel threats in cutting-edge AI/AGI systems. You will act as the operational arm of the security function, translating threat intelligence into real-time detections, coordinating incident response, and hunting for sophisticated attacks across GPU clusters and distributed training environments. This role is ideal for someone excited by the challenge of securing systems fundamentally different from traditional ones and defining the future of AI security engineering at scale.
Engineering Manager, GPU Infrastructure
Cohere
Cohere is a leading enterprise AI company building cutting-edge foundation AI models and end-to-end products. The GPU Clusters team is central to Cohere's infrastructure, responsible for building and operating the superclusters that power our frontier AI models. This role involves enabling research and development at the intersection of cutting-edge hardware, distributed systems, and AI research. As an Engineering Manager, you will lead a team of highly motivated engineers passionate about GPU infrastructure and AI, fostering a culture of technical excellence and innovation in a remote-first environment. This is a unique opportunity to shape the infrastructure powering the next generation of AI.
Software Engineer, Full Stack, Tinker
thinkingmachines
Thinking Machines Lab is seeking a full-stack engineer to develop and deploy the products and services that Tinker users engage with daily. This role involves working across frontend, backend, and infrastructure to build the Tinker console, developer tools, and other essential components for the platform. Tinker is a fine-tuning API that enables researchers and developers to customize frontier AI models using their own data and algorithms, managing the underlying infrastructure to provide flexibility and access to advanced capabilities.
$350k - $475k
Site Reliability Engineer (SRE)
thinkingmachines
Thinking Machines Lab is seeking a Site Reliability Engineer (SRE) to ensure the end-to-end reliability of their Tinker platform. This role involves working closely with engineers and research teams to enhance the robustness and resilience of every system layer. The SRE will be instrumental in maintaining and improving the infrastructure that supports custom AI model fine-tuning, ensuring a seamless experience for researchers and developers.
$350k - $475k
Research Engineer, Knowledge Foundations
Anthropic
The Knowledge Work team at Anthropic builds the training environments and evaluations that empower Claude to excel in professional workflows, including searching, analyzing, and creating content across various tools and documents. As this work expands, the underlying systems require the same rigor as the research itself. In this role, you will design and execute experiments to enhance Claude's ability to search, retrieve, and reason over information at scale. Your responsibilities will encompass environment design, data curation, RL training, evaluation, and supporting infrastructure, requiring flexibility to address progress blockers. You will collaborate closely with researchers and other RL teams to implement capabilities that directly influence Claude's performance. We believe that true ownership and impact stem from both hardening existing environments and creating new ones, ensuring the quality of the entire stack that drives superhuman epistemics.
Research Engineer, Pretraining
Anthropic
Anthropic is seeking a Research Engineer to join its Pretraining team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, aiming to build safe, steerable, and trustworthy AI systems. The mission is to ensure that transformative AI systems are aligned with human interests and are beneficial for society. The team is dedicated to pushing the boundaries of AI while prioritizing safety and ethics.
Research Engineer / Research Scientist, Pre-training
Anthropic
Anthropic is seeking passionate Research Scientists and Engineers to join our growing Pre-training team in Zurich. This role is at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems. You will be involved in developing the next generation of large language models, with a primary focus on multimodal capabilities, enabling LLMs to understand and interact with modalities beyond text.
Research Engineer/Research Scientist, Pre-training
Anthropic
Anthropic is seeking a Research Engineer to join its Pre-training team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, contributing to the creation of safe, steerable, and trustworthy AI systems. The team is dedicated to ensuring that transformative AI systems are aligned with human interests and societal benefit.
Staff Research Engineer, Discovery Team
Anthropic
Anthropic is dedicated to building reliable, interpretable, and steerable AI systems that are safe and beneficial for society. As a Research Engineer on the Discovery Team, you will work end-to-end to identify and address key blockers on the path to scientific Artificial General Intelligence (AGI). This role involves improving models' abilities to use computers, acting as a laboratory for long-horizon tasks and a crucial component for scientific workflows. You will collaborate with a team of researchers and engineers focused on pushing the scientific frontier.
Performance Engineer, GPU
Anthropic
Pioneering the next generation of AI requires breakthrough innovations in GPU performance and systems engineering. As a GPU Performance Engineer, you'll architect and implement the foundational systems that power Claude and push the frontiers of what's possible with large language models. You'll be responsible for maximizing GPU utilization and performance at unprecedented scale, developing cutting-edge optimizations that directly enable new model capabilities and dramatically improve inference efficiency. Working at the intersection of hardware and software, you'll implement state-of-the-art techniques from custom kernel development to distributed system architectures. Your work will span the entire stack—from low-level tensor core optimizations to orchestrating thousands of GPUs in perfect synchronization. Strong candidates will have a track record of delivering transformative GPU performance improvements in production ML systems and will be excited to shape the future of AI infrastructure alongside world-class researchers and engineers.
Research Engineer, Code RL (Reinforcement Learning)
Anthropic
We are seeking a Research Engineer for our Code RL team, focused on advancing AI models' capabilities in writing, editing, testing, debugging, and shipping real software. This role involves designing RL environments, coding tasks, and reward signals, as well as running training experiments on frontier models. You will diagnose model performance, improve pipeline speed and reliability, and contribute to areas like agentic coding behaviors, code correctness, and autonomous engineering. The position blends cutting-edge research with practical engineering to build high-quality, scalable AI systems.
Research Scientist, Data
pika
Pika is seeking a staff or lead-level Research Engineer, Data to architect and scale data engineering systems for their advanced multimodal foundation models. This role is crucial for strengthening research teams by building, optimizing, and owning large-scale data pipelines and ML data curation. The goal is to ensure foundation models have access to high-quality, diverse datasets, enabling millions of creators. If you are passionate about data infrastructure and innovative research-engineering, this is an opportunity to make a significant impact.
Research Engineer, Post-Training
Harvey
Harvey is transforming legal and professional services by combining agentic AI, an enterprise-grade platform, and deep domain expertise. We are seeking a Research Engineer focused on post-training to scale the process of turning expert feedback and agent traces into significantly improved models. This role involves defining and running model training experiments, interpreting results, and collaborating with internal and external partners to enhance data, environments, graders, and training methodologies. The ideal candidate is a self-manager with extensive hands-on experience training open-weight models and the engineering depth to execute and debug experiments efficiently.
$231k - $340k
ML Engineer, Inference & Optimization
pika
We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale. You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.
Research Engineer, Frontier Speculative Decoding
Together AI
Together AI is building the Inference Platform that powers the world's most advanced generative AI models. This role will serve as a critical bridge between cutting-edge research and real-world applications, focusing on translating internal model training research into production-ready deployments for customers. The work involves a deep commitment to data-centric development, meticulous hyperparameter tuning, and rigorous checkpoint evaluation. You will transform general-purpose models into highly performant, specialized tools by fine-tuning them on customer-specific data and internal datasets, working with dedicated GPU clusters rather than training foundation models from scratch.
$190k - $270k
Research Intern, Model Shaping (Fall 2026)
Together AI
As a Research Intern in the Model Shaping team, you will work on advanced post-training methods, new techniques for efficient neural network training, and robust evaluation of foundation model capabilities. The Model Shaping team at Together AI focuses on tailoring open foundation models for downstream applications, building services for machine learning developers, and developing new methods for efficient model training and evaluation. This role offers the opportunity to contribute to cutting-edge research and potentially influence open-source projects.
Technical Lead Manager, Physical AI
Scale AI
Scale AI is seeking a Technical Lead Manager for its Physical AI team, focusing on the development of general AI that can reason and act in the physical world. This role bridges cutting-edge Machine Learning research with physical robot deployment, leading a team of Research Engineers while remaining a hands-on technical contributor. The primary focus is on developing and evaluating Large-Scale Foundation Models, such as VLAs and World models, to enable robots and autonomous vehicles to generalize across diverse tasks and environments. The team leverages Scale's extensive data infrastructure to help build Foundation Models for Physical AI, aiming to redefine the future of automation.
$249k - $311k
Security engineer, detection and response
Writer
WRITER is seeking a Staff Detection and Response Engineer to join its security team. This role is crucial for protecting the AI infrastructure that is transforming how the world works. You will be responsible for building sophisticated detection systems to identify attacks targeting our AI platform, training data, and model deployments, while also creating automated response capabilities that scale with the company's rapid growth. This position offers a unique opportunity to combine hands-on security engineering with strategic thinking to stay ahead of novel threats in cutting-edge AI/AGI systems. You will act as the operational arm of the security function, translating threat intelligence into real-time detections, coordinating incident response, and hunting for sophisticated attacks across GPU clusters and distributed training environments. This role is ideal for someone excited by the challenge of securing systems fundamentally different from traditional ones and defining the future of AI security engineering at scale.
Security engineer, detection and response (UK)
Writer
WRITER is seeking a Staff Detection and Response Engineer to join its security team and protect the company's AI infrastructure. This role involves building sophisticated detection systems for attacks targeting the AI platform, training data, and model deployments, as well as creating automated response capabilities. The position combines hands-on security engineering with strategic thinking to address novel threats in cutting-edge AI/AGI systems. The engineer will be the operational arm of the security function, translating threat intelligence into real-time detections, coordinating incident response, and hunting for sophisticated attacks across GPU clusters and distributed training environments. This is an opportunity to define AI security engineering at scale, working closely with AI Security research, Cloud Infrastructure, Software Security Engineering, and AI researchers.