Distributed Training Jobs
55 open roles mentioning Distributed Training
Senior Member of Technical Staff, Safety and Security for Agents
Cohere
Cohere is seeking a Senior Member of Technical Staff to join the Safety and Security for Agents team. In this role, you will significantly contribute to the development of safer, fairer, more trustworthy, and more secure Large Language Models (LLMs). Your work will focus on data generation, post-training algorithms, and evaluation methods to ensure the safety of next-generation models that interact with external resources and take actions. You will collaborate closely with machine learning teams, data annotation teams, and product and policy teams, requiring a blend of machine learning expertise, ethical AI principles, experimental design, and data management skills. This position offers significant autonomy and decision-making power within a small team, with the opportunity to shape the future of LLMs for societal benefit.
Staff Cloud Support Engineer
crusoe
As a Staff Cloud Support Engineer at Crusoe, you will be a technical authority within Crusoe Cloud, acting as a force multiplier for Customer Experience, SRE, Networking, Fleet, and Product teams. Your role extends beyond simple ticket resolution; you will design reliability guardrails, influence architectural decisions, mentor other engineers, and directly contribute to revenue protection by preventing large-scale incidents. This position requires deep expertise in Linux systems, Kubernetes, networking, and AI/ML infrastructure, applied with a strong customer focus. You should be comfortable operating in ambiguous environments, leading incident response efforts, and shaping the scalability of Crusoe's high-performance AI infrastructure globally.
$156k - $190k
Research Engineer, Machine Learning
Mistral AI
Mistral AI is democratizing AI through high-performance, optimized, open-source models and solutions. We are a dynamic, collaborative team passionate about AI's potential to transform society, with a diverse workforce driving innovation. As a Research Engineer – ML track, you will build and optimize large-scale learning systems powering our open-weight models. You will work hand-in-hand with Research Scientists, either enhancing the shared training framework and data pipelines or embedding within a research squad to turn fresh ideas into scalable code.
Software Engineer - Training Product
Baseten
Baseten is seeking a customer-obsessed software engineer to join their team and contribute to the development of mission-critical AI inference platforms. In this role, you will own features from conception to launch, working across the entire technology stack from API and UI down to the infrastructure layer. You will have the opportunity to fine-tune models, gain a deep understanding of user workflows, and collaborate closely with research engineers to build cutting-edge experiences that accelerate model development and address real-world pain points. If you are excited about diving deep into AI model training and building impactful products, this is the role for you.
Member of Technical Staff, Data Analysis and Evaluation
Cohere
Cohere is seeking a Member of Technical Staff in Data Analysis and Evaluation to ensure the quality, reliability, and performance of our large language models (LLMs). This role involves designing and conducting data collection tasks, assessing dataset quality, and analyzing model robustness and generalisability. You will collaborate with researchers, engineers, and data annotators to drive data-driven decisions and enhance AI system effectiveness. The position requires expertise in statistics, experimental design, and machine learning to ensure high-quality data and reliable model performance across diverse scenarios, contributing to Cohere's mission of advancing AI.
Senior ML Systems Engineer, Frameworks & Tooling
Cohere
Cohere is seeking a Senior ML Systems Engineer to join their team and build, maintain, and evolve the training framework that powers their frontier-scale language models. This role is ideal for someone passionate about large-scale training, distributed systems, and HPC infrastructure, offering the opportunity to design and maintain core components for fast, reliable, and scalable model training. You will also build tooling to connect research ideas to thousands of GPUs, working across the full stack of ML systems with significant autonomy and impact.
Member of Technical Staff, LLM Infrastructure
fireworks ai
As a Software Engineer on the AI Infrastructure team, you will help design the core systems that power Fireworks AI’s generative AI platform. You will build infrastructure and tools that ensure the reliability, performance, quality, and availability of our AI system. Your mission is to make Fireworks AI the most reliable and user-friendly generative AI platform in the world. You will partner closely with our cloud infrastructure, product, and performance teams to deliver infrastructure that bridges the gap between our customers and the ultra-performant proprietary Fireworks inference engine.
Forward Deployed Engineer, Lead - LLM Post-training
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible. We are seeking an exceptional technical leader to build and scale our post-training and evaluation capabilities within the Applied AI team. This role involves taking our open-weight models and adapting them for specific customer domains, tasks, and constraints. You will own the end-to-end technical strategy for model customization, from synthetic data generation and reward modeling through training and production deployment, working directly with customers and research teams.
Member of Technical Staff - Pre-Training
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible for everyone. We build open models that empower individuals to control their intelligence and shape the future of AI. As a Member of Technical Staff - Pre-Training, you will be instrumental in researching and building solutions across algorithms, scaling laws, data processing, optimizers, and model architecture. This role involves designing and executing scientific experiments to deepen our understanding of scaling large language models and data efficiency, while also implementing state-of-the-art deep learning methods. You will have the opportunity to lead small research projects independently and contribute to larger initiatives, optimizing training infrastructure for efficient scaling and working across the entire stack from low-level optimizations to high-level model design.
Applied Machine Learning Engineer
Cohere
Cohere is seeking a Member of Technical Staff for their Applied ML team. In this role, you will collaborate directly with customers to understand their challenges and implement solutions leveraging Large Language Models. You will apply your problem-solving skills, creativity, and technical expertise to bridge the gap in enterprise AI adoption, delivering impactful products and disrupting key industries. This is an opportunity to join at a pivotal moment, shape the company's offerings, and contribute to cutting-edge AI development.
Machine Learning Infrastructure Engineer, Model Inference
Abridge
Abridge is seeking an ML Infrastructure Engineer, Model Inference to build and optimize the core inference infrastructure powering their machine learning models. This role is crucial for enhancing the scalability, efficiency, and performance of Abridge's AI-driven healthcare solutions. The engineer will collaborate with Infrastructure and Research teams to build, deploy, optimize, and orchestrate AI models, working on a platform that transforms patient-clinician conversations into structured clinical notes in real-time.
Member of Technical Staff, Integration/RL Team (Research Engineer)
Cohere
Cohere is a leading enterprise AI company focused on building cutting-edge foundation AI models and end-to-end products for real-world business problems. The integration team specifically focuses on developing and scaling machine learning algorithms and infrastructure for LLM post-training, with an emphasis on large-scale, distributed Reinforcement Learning (RL) methods. This role is crucial for enhancing the post-training codebase by implementing new research tools, optimizing algorithms, and scaling distributed RL capabilities. We are looking for passionate individuals who are meticulous in their approach to engineering and science, contributing to both production code and research efforts.
Member of Technical Staff, Post-Training
Cohere
Cohere is seeking a Member of Technical Staff to focus on post-training of AI models. This role is crucial for advancing the state of the art in model post-training and shipping cutting-edge models to production, bridging the gap between research and practical application. You will have access to significant compute resources and a talented team to contribute to increasing model capabilities and driving customer value. The position offers a unique opportunity to contribute to both production code and research efforts, depending on your interests and organizational needs.
Member of Technical Staff - Image / Video Generation
Black Forest Labs
We are seeking a Member of Technical Staff specializing in Image/Video Generation to join our team. You will be instrumental in training large-scale diffusion models for image and video generation, pushing the boundaries of generative AI. This role involves exploring novel approaches, rigorously testing design choices, and understanding the trade-offs between speed and quality in production settings. You will contribute to shaping our research direction by conducting experiments, analyzing results, and communicating findings to the team. This is an opportunity to work on foundational technologies used by millions of creators worldwide and contribute to the advancement of generative models.
€130k - €340k
Research Engineer, Machine Learning - Paris/London/Zurich/Warsaw
Mistral AI
Mistral AI is a pioneering company focused on democratizing AI through high-performance, optimized, open-source models and solutions. We aim to simplify tasks, save time, and enhance learning and creativity by integrating AI seamlessly into daily working life. Our comprehensive AI platform serves both enterprise and personal needs, featuring offerings like Le Chat, La Plateforme, Mistral Code, and Mistral Compute. We are a dynamic, collaborative, and diverse team passionate about AI's potential to transform society, driven by innovation and a low-ego, team-spirited culture.