Together AI

Open positions (32)

Systems Research Engineer Intern - GPU Programming (Fall 2026)

1mo ago
Together AI

Together AI

As a Systems Research Engineer Intern specialized in GPU Programming, you will play a crucial role in developing and optimizing GPU-accelerated kernels and algorithms for ML/AI applications. You will co-design GPU kernels and model architecture with the modeling and algorithm team to enhance the performance and efficiency of our AI systems. Collaborating with the hardware and software teams, you will contribute to the co-design of efficient GPU architectures and programming models, leveraging your expertise in GPU programming and parallel computing. Your research skills will be vital in staying up-to-date with the latest advancements in GPU programming techniques, ensuring that our AI infrastructure remains at the forefront of innovation.

San Francisco remote
CUDATritonParallel Computing +4 more

Systems Research Engineer, GPU Programming

1mo ago
Together AI

Together AI

As a Systems Research Engineer specialized in GPU Programming, you will play a crucial role in developing and optimizing GPU-accelerated kernels and algorithms for ML/AI applications. You will co-design GPU kernels and model architecture to enhance the performance and efficiency of our AI systems, and contribute to the co-design of efficient GPU architectures and programming models. Your research skills will be vital in staying up-to-date with the latest advancements in GPU programming techniques, ensuring that our AI infrastructure remains at the forefront of innovation.

$160k - $230k

San Francisco remote
AIMLCUDA +5 more

Staff Machine Learning Engineer, Voice AI

1mo ago
Together AI

Together AI

Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Staff ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro, focusing on pushing latency and throughput boundaries. You will address unique challenges in voice inference, such as streaming audio and real-time latency, and shape the future of how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.

$220k - $280k

San Francisco remote
PythonGoFine-Tuning +10 more

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

1mo ago
Together AI

Together AI

Together AI is seeking a Staff Engineer to design and deliver multi-petabyte storage systems optimized for large-scale AI training and inference workloads. You will architect high-performance parallel filesystems and object stores, integrate cutting-edge technologies, and drive significant cost optimization. The role involves building Kubernetes-native storage operators and self-service platforms for automated provisioning and multi-tenancy. You will focus on optimizing data paths, designing multi-tier caching architectures, and tuning parallel filesystems for AI applications. This is a research-driven role within a company focused on lowering the cost of modern AI systems through co-design of software, hardware, algorithms, and models.

$250k - $300k

San Francisco remote
KubernetesPythonGo +7 more

Senior Machine Learning Engineer, Voice AI

1mo ago
Together AI

Together AI

Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Senior ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro to achieve frontier-level latency and throughput. You will focus on unique voice inference challenges such as streaming audio, tokenization, and real-time latency budgets, shaping how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.

$200k - $260k

San Francisco remote
PythonGoFine-Tuning +10 more

Research Engineer, Frontier Speculative Decoding

1mo ago
Together AI

Together AI

Together AI is building the Inference Platform that powers the world's most advanced generative AI models. This role will serve as a critical bridge between cutting-edge research and real-world applications, focusing on translating internal model training research into production-ready deployments for customers. The work involves a deep commitment to data-centric development, meticulous hyperparameter tuning, and rigorous checkpoint evaluation. You will transform general-purpose models into highly performant, specialized tools by fine-tuning them on customer-specific data and internal datasets, working with dedicated GPU clusters rather than training foundation models from scratch.

$190k - $270k

San Francisco, New York City remote
KubernetesPythonFine-Tuning +6 more

Research Engineer, Core ML

1mo ago
Together AI

Together AI

This research engineering role focuses on translating new Reinforcement Learning (RL) algorithms, scheduling methods, and inference optimizations into production-grade systems that power Together's API. The Core ML team operates at the intersection of efficient inference (algorithms, architectures, engines) and post-training/RL systems, building and maintaining high-performance inference and RL engines at production scale. The goal is to significantly improve model speed, cost-efficiency, and capabilities through RL-based post-training. This position requires a blend of algorithmic understanding and systems engineering, with opportunities to work across the entire stack from RL algorithms and training engines to kernels and serving systems, ultimately driving measurable improvements in latency, throughput, cost, and model quality at scale.

$200k - $280k

San Francisco remote
PythonTransformersRLHF +8 more

Machine Learning, Platform Engineer

1mo ago
Together AI

Together AI

Together AI is a research-driven artificial intelligence company focused on lowering the cost of modern AI systems. This role is part of a team dedicated to enabling custom models and dedicated inference on Together's platform. The team is responsible for building a container platform, optimizing autoscaling, minimizing cold starts, achieving the best end-to-end model performance, and providing a best-in-class developer experience with great tooling. The work often involves video or audio generation across the stack, including CUDA kernels, PyTorch optimization, inference engines, container orchestration, and queueing theory.

$160k - $250k

San Francisco remote
KubernetesPythonRust +7 more

Machine Learning Engineer - Inference

1mo ago
Together AI

Together AI

Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large language models and ensuring they run efficiently and effectively at scale. You will collaborate closely with AI researchers and engineers to create cutting-edge AI solutions and shape the future of AI inference.

$160k - $230k

San Francisco remote
PythonRustPyTorch +6 more

Machine Learning Engineer

1mo ago
Together AI

Together AI

Together AI is seeking an ML Engineer to develop systems and APIs for customer inference and fine-tuning of LLMs. The ideal candidate will have experience implementing runtime systems for large-scale AI/ML model inference, including the largest LLMs. This role involves designing and building production systems for reliability and performance at scale, partnering with cross-functional teams, and improving system efficiency and stability.

$160k - $220k

San Francisco remote
PythonGoRust +7 more

LLM Inference Frameworks and Optimization Engineer

1mo ago
Together AI

Together AI

Together.ai is building state-of-the-art infrastructure for efficient and scalable inference of large language models (LLMs). The company's mission is to optimize inference frameworks, algorithms, and infrastructure to push the boundaries of performance, scalability, and cost-efficiency. They are seeking an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines for multimodal and language models at scale. This role will focus on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design, ensuring efficient large-scale deployment of LLMs and vision models. This position offers a unique opportunity to shape the future of LLM inference infrastructure and ensure scalable, high-performance AI deployment across diverse applications.

$160k - $230k

San Francisco, Singapore, Amsterdam remote
KubernetesPythonC# +9 more

AI infrastructure Engineer (SRE) Amsterdam

1mo ago
Together AI

Together AI

Together AI is seeking an AI Infrastructure Engineer (SRE) to ensure the smooth operation of user-facing services and production systems. This role combines the skills of a pragmatic operator and a software engineer, applying sound engineering principles, operational discipline, and automation to our operating environments and codebase. You will specialize in systems such as operating systems, storage subsystems, and networking, while implementing best practices for availability, reliability, and scalability, with interests in algorithms and distributed systems. Join a research-driven artificial intelligence company focused on advancing AI through open and transparent systems, aiming to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models.

Amsterdam onsite
KubernetesTerraformObservability +7 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.