vLLM Jobs

7 open roles mentioning vLLM

Machine Learning Engineer

1mo ago
Together AI

Together AI

Together AI is seeking an ML Engineer to develop systems and APIs for customer inference and fine-tuning of LLMs. The ideal candidate will have experience implementing runtime systems for large-scale AI/ML model inference, including the largest LLMs. This role involves designing and building production systems for reliability and performance at scale, partnering with cross-functional teams, and improving system efficiency and stability.

$160k - $220k

San Francisco remote
PythonGoRust +7 more

Senior Machine Learning Engineer, Voice AI

1mo ago
Together AI

Together AI

Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Senior ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro to achieve frontier-level latency and throughput. You will focus on unique voice inference challenges such as streaming audio, tokenization, and real-time latency budgets, shaping how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.

$200k - $260k

San Francisco remote
PythonGoFine-Tuning +10 more

Staff Machine Learning Engineer, Voice AI

1mo ago
Together AI

Together AI

Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Staff ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro, focusing on pushing latency and throughput boundaries. You will address unique challenges in voice inference, such as streaming audio and real-time latency, and shape the future of how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.

$220k - $280k

San Francisco remote
PythonGoFine-Tuning +10 more

AI Researcher, Core ML (Turbo)

1mo ago
Together AI

Together AI

The Turbo team operates at the intersection of efficient inference (algorithms, architectures, engines) and post-training/RL systems. We are responsible for building and managing the systems that power Together's API, focusing on high-performance inference and RL/post-training engines capable of operating at production scale. Our core mission is to advance the frontiers of efficient inference and RL-driven training, aiming to make models significantly faster and more cost-effective to run, while simultaneously enhancing their capabilities through RL-based post-training methods. This role involves working across the entire stack, from RL algorithms and training engines to kernels and serving systems, to develop and refine state-of-the-art models using RL pipelines. We value individuals with deep expertise in one area and a strong willingness to collaborate and grow across others.

$200k - $280k

San Francisco remote
PythonTransformersRLHF +7 more

Forward Deployed Engineer (Inference & Post-Training)

1mo ago
Together AI

Together AI

As a Forward Deployed Engineer (FDE) focused on Inference & Post-Training, you will be a hands-on technical partner to strategic customers, assisting production AI teams with leveraging high-quality models and performing inference at scale. You will act as a deep-domain specialist in inference optimization, fine-tuning pipelines, and production deployment, partnering with Solutions Architects. FDEs add significant value by ensuring complex Proofs of Concept (POCs) are met, facilitating platform adoption, and guiding tailored optimization efforts, directly impacting customer success and company growth.

$270k - $300k

San Francisco remote
PythonFine-TuningRLHF +9 more

Senior Backend Engineer, Inference Platform

1mo ago
Together AI

Together AI

Together AI is building the Inference Platform to bring advanced generative AI models to the world, powering multi-tenant serverless workloads and dedicated endpoints. This role offers a unique opportunity to optimize latency and fully utilize tens of thousands of GPUs, working hands-on with cutting-edge hardware. You will collaborate directly with research teams to productionize frontier models and engage with the open-source community, contributing to projects that push the boundaries of inference performance and efficiency.

$160k - $250k

San Francisco remote
KubernetesPythonTypeScript +7 more

Open-Source Software, Machine Learning Engineer

1y ago
Mistral AI

Mistral AI

Mistral AI is democratizing AI through high-performance, optimized, open-source models, products, and solutions. We are a dynamic, collaborative team passionate about AI's potential to transform society, with teams distributed globally. We are seeking an Open-Source Software, Machine Learning Engineer to join our OSS team, which is embedded within our Science team. This role is critical in helping turn research breakthroughs into tangible solutions and improving Mistral's open-source ecosystem by open-sourcing state-of-the-art models and maintaining our publicly available libraries.

Paris remote Full-time
MistralPythonGo +10 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.