CUDA Jobs

16 open roles mentioning CUDA

Research Intern RL & Post-Training Systems, Turbo (Fall 2026)

1mo ago
Together AI

Together AI

The Turbo Research team focuses on making post-training and reinforcement learning for large language models efficient, scalable, and reliable. This work intersects RL algorithms, inference systems, and large-scale experimentation, where inference costs significantly impact training efficiency and the practicality of learning algorithms. As a research intern, you will investigate RL and post-training methods whose performance and scalability are closely tied to inference behavior, co-designing algorithms and systems. Projects aim to enable new experimental regimes, including larger models, longer rollouts, and more complex evaluations, by re-evaluating the interaction between inference, scheduling, and training.

San Francisco remote
PythonC#NLP +7 more

LLM Inference Frameworks and Optimization Engineer

1mo ago
Together AI

Together AI

Together.ai is building state-of-the-art infrastructure for efficient and scalable inference of large language models (LLMs). The company's mission is to optimize inference frameworks, algorithms, and infrastructure to push the boundaries of performance, scalability, and cost-efficiency. They are seeking an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines for multimodal and language models at scale. This role will focus on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design, ensuring efficient large-scale deployment of LLMs and vision models. This position offers a unique opportunity to shape the future of LLM inference infrastructure and ensure scalable, high-performance AI deployment across diverse applications.

$160k - $230k

San Francisco, Singapore, Amsterdam remote
KubernetesPythonC# +9 more

Machine Learning Engineer - Inference

1mo ago
Together AI

Together AI

Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large language models and ensuring they run efficiently and effectively at scale. You will collaborate closely with AI researchers and engineers to create cutting-edge AI solutions and shape the future of AI inference.

$160k - $230k

San Francisco remote
PythonRustPyTorch +6 more

Machine Learning, Platform Engineer

1mo ago
Together AI

Together AI

Together AI is a research-driven artificial intelligence company focused on lowering the cost of modern AI systems. This role is part of a team dedicated to enabling custom models and dedicated inference on Together's platform. The team is responsible for building a container platform, optimizing autoscaling, minimizing cold starts, achieving the best end-to-end model performance, and providing a best-in-class developer experience with great tooling. The work often involves video or audio generation across the stack, including CUDA kernels, PyTorch optimization, inference engines, container orchestration, and queueing theory.

$160k - $250k

San Francisco remote
KubernetesPythonRust +7 more

Senior Machine Learning Engineer, Voice AI

1mo ago
Together AI

Together AI

Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Senior ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro to achieve frontier-level latency and throughput. You will focus on unique voice inference challenges such as streaming audio, tokenization, and real-time latency budgets, shaping how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.

$200k - $260k

San Francisco remote
PythonGoFine-Tuning +10 more

Staff Machine Learning Engineer, Voice AI

1mo ago
Together AI

Together AI

Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Staff ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro, focusing on pushing latency and throughput boundaries. You will address unique challenges in voice inference, such as streaming audio and real-time latency, and shape the future of how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.

$220k - $280k

San Francisco remote
PythonGoFine-Tuning +10 more

Systems Research Engineer, GPU Programming

1mo ago
Together AI

Together AI

As a Systems Research Engineer specialized in GPU Programming, you will play a crucial role in developing and optimizing GPU-accelerated kernels and algorithms for ML/AI applications. You will co-design GPU kernels and model architecture to enhance the performance and efficiency of our AI systems, and contribute to the co-design of efficient GPU architectures and programming models. Your research skills will be vital in staying up-to-date with the latest advancements in GPU programming techniques, ensuring that our AI infrastructure remains at the forefront of innovation.

$160k - $230k

San Francisco remote
AIMLCUDA +5 more

Systems Research Engineer Intern - GPU Programming (Fall 2026)

1mo ago
Together AI

Together AI

As a Systems Research Engineer Intern specialized in GPU Programming, you will play a crucial role in developing and optimizing GPU-accelerated kernels and algorithms for ML/AI applications. You will co-design GPU kernels and model architecture with the modeling and algorithm team to enhance the performance and efficiency of our AI systems. Collaborating with the hardware and software teams, you will contribute to the co-design of efficient GPU architectures and programming models, leveraging your expertise in GPU programming and parallel computing. Your research skills will be vital in staying up-to-date with the latest advancements in GPU programming techniques, ensuring that our AI infrastructure remains at the forefront of innovation.

San Francisco remote
CUDATritonParallel Computing +4 more

Senior Backend Engineer, Inference Platform

1mo ago
Together AI

Together AI

Together AI is building the Inference Platform to bring advanced generative AI models to the world, powering multi-tenant serverless workloads and dedicated endpoints. This role offers a unique opportunity to optimize latency and fully utilize tens of thousands of GPUs, working hands-on with cutting-edge hardware. You will collaborate directly with research teams to productionize frontier models and engage with the open-source community, contributing to projects that push the boundaries of inference performance and efficiency.

$160k - $250k

San Francisco remote
KubernetesPythonTypeScript +7 more

ML Systems Engineer, Robotics

2mo ago
Scale AI

Scale AI

Scale's Physical AI business unit is focused on solving data bottlenecks in Robotics, Autonomous Vehicles, and Computer Vision. This role involves applied research and developing ML pipelines for processing, training, and fine-tuning data collected by Scale, with an emphasis on optimizing algorithms and pipelines for efficient GPU execution in the cloud. You will advance research, shape Scale's offerings, and expand the frontier of data and model evaluation for Physical AI. As an ML Systems Engineer, you will design and build platforms for scalable, reliable, and efficient serving of foundation models tailored for physical agents, powering both internal research and external customer use cases.

$249k - $311k

San Francisco, CA remote
AWSDockerKubernetes +7 more

Machine Learning Research Engineer, Agent Data Foundation - Enterprise GenAI

2mo ago
Scale AI

Scale AI

Scale is seeking a Machine Learning Research Engineer to join the Agent Data Foundation team within the Enterprise ML Research Lab. This role will focus on defining and building the data flywheel that drives AI development, particularly for complex agents in enterprise settings. You will conduct cutting-edge research in areas such as synthetic data generation for RL post-training, building agents for production trace analysis, and contributing to a framework for agent construction. The goal is to create state-of-the-art agents through advanced post-training and agent-building algorithms, shaping the future of the GenAI movement.

$265k - $331k

San Francisco, CA; New York, NY onsite
PythonPyTorchLLMs +4 more

Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI

2mo ago
Scale AI

Scale AI

Scale is seeking a Machine Learning Systems Research Engineer to join their Enterprise ML Research Lab. This role will focus on building algorithms for a next-generation Agent RL training platform, supporting large-scale training, and integrating state-of-the-art technologies to optimize ML systems. You will collaborate with other ML researchers and engineers who apply these algorithms to client use cases, including AI cybersecurity firewalls and healthtech search models. If you are passionate about shaping the future of AI, this is an exciting opportunity to contribute to cutting-edge advancements in enterprise GenAI.

$265k - $331k

San Francisco, CA; New York, NY onsite
LLMPyTorchTransformers +7 more

Tech Lead Manager- MLRE, ML Systems

2mo ago
Scale AI

Scale AI

Scale's LLM post-training platform team builds our internal distributed framework for large language model training, powering MLEs, researchers, data scientists, and operators for fast and automatic training and evaluation of LLMs. This platform also serves as the underlying training framework for the data quality evaluation pipeline. You will work closely with Scale’s ML teams and researchers to build the foundation platform which supports all our ML research and development works, optimizing it to enable next generation LLM training, inference, and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you!

$265k - $331k

San Francisco, CA; New York, NY onsite
LLMPyTorchTransformers +7 more

ML Research Engineer, ML Systems

2mo ago
Scale AI

Scale AI

Scale's ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. This platform powers MLEs, researchers, data scientists, and operators for fast and automatic training and evaluation of LLMs, as well as data quality evaluation. You will work closely across Scale’s ML teams and researchers to build the foundation platform that supports all our ML research and development, optimizing it to enable the next generation of LLM training, inference, and data curation. If you are excited about shaping the future of AI via fundamental innovations, we would love to hear from you!

$190k - $237k

San Francisco, CA; Seattle, WA; New York, NY onsite
LLMPyTorchTransformers +7 more

Research Scientist – Controlled 3D Generation

4mo ago
Stability AI

Stability AI

We are seeking a Research Scientist passionate about 3D generation, flow matching, and diffusion models. You will help advance the frontier of controllable 3D content creation by building models that generate consistent, editable, and physically grounded 3D assets and scenes. This role involves conducting cutting-edge research, designing and implementing scalable training pipelines, and developing techniques for conditioning and control. You will analyze model behavior, collaborate with cross-disciplinary teams to translate research into production-ready systems, and publish results at top-tier venues.

Remote remote
PyTorchJAXCUDA +7 more

Research Engineer, Machine Learning - Paris/London/Zurich/Warsaw

1y ago
Mistral AI

Mistral AI

Mistral AI is a pioneering company focused on democratizing AI through high-performance, optimized, open-source models and solutions. We aim to simplify tasks, save time, and enhance learning and creativity by integrating AI seamlessly into daily working life. Our comprehensive AI platform serves both enterprise and personal needs, featuring offerings like Le Chat, La Plateforme, Mistral Code, and Mistral Compute. We are a dynamic, collaborative, and diverse team passionate about AI's potential to transform society, driven by innovation and a low-ego, team-spirited culture.

Paris hybrid Full-time
MistralPythonPyTorch +11 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.