vLLM Jobs
7 open roles mentioning vLLM
Machine Learning Engineer
Together AI
Together AI is seeking an ML Engineer to develop systems and APIs for customer inference and fine-tuning of LLMs. The ideal candidate will have experience implementing runtime systems for large-scale AI/ML model inference, including the largest LLMs. This role involves designing and building production systems for reliability and performance at scale, partnering with cross-functional teams, and improving system efficiency and stability.
$160k - $220k
Senior Machine Learning Engineer, Voice AI
Together AI
Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Senior ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro to achieve frontier-level latency and throughput. You will focus on unique voice inference challenges such as streaming audio, tokenization, and real-time latency budgets, shaping how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.
$200k - $260k
Staff Machine Learning Engineer, Voice AI
Together AI
Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Staff ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro, focusing on pushing latency and throughput boundaries. You will address unique challenges in voice inference, such as streaming audio and real-time latency, and shape the future of how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.
$220k - $280k
AI Researcher, Core ML (Turbo)
Together AI
The Turbo team operates at the intersection of efficient inference (algorithms, architectures, engines) and post-training/RL systems. We are responsible for building and managing the systems that power Together's API, focusing on high-performance inference and RL/post-training engines capable of operating at production scale. Our core mission is to advance the frontiers of efficient inference and RL-driven training, aiming to make models significantly faster and more cost-effective to run, while simultaneously enhancing their capabilities through RL-based post-training methods. This role involves working across the entire stack, from RL algorithms and training engines to kernels and serving systems, to develop and refine state-of-the-art models using RL pipelines. We value individuals with deep expertise in one area and a strong willingness to collaborate and grow across others.
$200k - $280k
Forward Deployed Engineer (Inference & Post-Training)
Together AI
As a Forward Deployed Engineer (FDE) focused on Inference & Post-Training, you will be a hands-on technical partner to strategic customers, assisting production AI teams with leveraging high-quality models and performing inference at scale. You will act as a deep-domain specialist in inference optimization, fine-tuning pipelines, and production deployment, partnering with Solutions Architects. FDEs add significant value by ensuring complex Proofs of Concept (POCs) are met, facilitating platform adoption, and guiding tailored optimization efforts, directly impacting customer success and company growth.
$270k - $300k
Senior Backend Engineer, Inference Platform
Together AI
Together AI is building the Inference Platform to bring advanced generative AI models to the world, powering multi-tenant serverless workloads and dedicated endpoints. This role offers a unique opportunity to optimize latency and fully utilize tens of thousands of GPUs, working hands-on with cutting-edge hardware. You will collaborate directly with research teams to productionize frontier models and engage with the open-source community, contributing to projects that push the boundaries of inference performance and efficiency.
$160k - $250k
Open-Source Software, Machine Learning Engineer
Mistral AI
Mistral AI is democratizing AI through high-performance, optimized, open-source models, products, and solutions. We are a dynamic, collaborative team passionate about AI's potential to transform society, with teams distributed globally. We are seeking an Open-Source Software, Machine Learning Engineer to join our OSS team, which is embedded within our Science team. This role is critical in helping turn research breakthroughs into tangible solutions and improving Mistral's open-source ecosystem by open-sourcing state-of-the-art models and maintaining our publicly available libraries.