Model Serving Jobs
16 open roles mentioning Model Serving
Staff Software Engineer, AI Reliability
Anthropic
Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. The AI Reliability Engineering (AIRE) team partners with other teams across Anthropic to enhance the reliability of critical serving paths, from SDKs through API layers, serving infrastructure, and accelerators. This role offers a unique, cross-cutting exposure to the most important systems at Anthropic, requiring a holistic view of system composition and reliability.
Staff Software Engineer, AI Reliability Engineering
Anthropic
AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths, from the SDK through our network, API layers, serving infrastructure, and accelerators. This role involves jumping into the trenches alongside partner teams to make the systems that deliver Claude more robust and resilient, be it during an incident or collaborating on projects. Reliability here is an emergent phenomenon that transcends any single team's boundaries, requiring someone to zoom out and look at the whole picture, offering dynamic, cross-cutting exposure to the systems that matter most.
Staff Software Engineer, AI Reliability Engineering
Anthropic
AI Reliability Engineering (AIRE) partners with teams across Anthropic to improve reliability across our most critical serving paths, from SDKs through our network, API layers, serving infrastructure, and accelerators. This role involves jumping into the trenches alongside partner teams to make the systems that deliver Claude more robust and resilient, whether during an incident or collaborating on projects. Reliability is viewed as an emergent phenomenon that transcends single team boundaries, requiring a holistic perspective across the entire system.
Machine Learning Engineer, Reliability
Fal
fal is building the generative media ecosystem for the next generation of AI products, providing the infrastructure, tools, and model access needed to scale from idea to production. As generative media reshapes industries, fal is becoming the foundation for ambitious teams. This hybrid ML Engineering / Site Reliability Engineering role will own the reliability, security, and safety of fal's generative media model APIs, ensuring they remain available, performant, secure, and safe for thousands of developers and enterprises. You will address model-specific failure modes, such as degraded output quality, drift, unsafe generations, and abuse patterns, as critical reliability concerns alongside uptime and latency.
ML Engineer, Inference & Optimization
pika
We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale. You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.
LLM Inference Frameworks and Optimization Engineer
Together AI
Together.ai is building state-of-the-art infrastructure for efficient and scalable inference of large language models (LLMs). The company's mission is to optimize inference frameworks, algorithms, and infrastructure to push the boundaries of performance, scalability, and cost-efficiency. They are seeking an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines for multimodal and language models at scale. This role will focus on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design, ensuring efficient large-scale deployment of LLMs and vision models. This position offers a unique opportunity to shape the future of LLM inference infrastructure and ensure scalable, high-performance AI deployment across diverse applications.
$160k - $230k
Senior Machine Learning Engineer, Voice AI
Together AI
Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Senior ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro to achieve frontier-level latency and throughput. You will focus on unique voice inference challenges such as streaming audio, tokenization, and real-time latency budgets, shaping how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.
$200k - $260k
Staff Machine Learning Engineer, Voice AI
Together AI
Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Staff ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro, focusing on pushing latency and throughput boundaries. You will address unique challenges in voice inference, such as streaming audio and real-time latency, and shape the future of how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.
$220k - $280k
Senior Platform Engineer, Voice AI
Together AI
Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications with best-in-class latency and reliability. We are seeking a Senior Platform Engineer to take ownership of the API and infrastructure layer for voice workloads. You will develop the real-time WebSocket and HTTP APIs used by developers to deploy voice experiences, design autoscaling for latency-sensitive streaming workloads, and ensure the reliability of our multi-provider voice platform for production voice agents handling millions of calls. This is a critical, foundational role on a small, high-impact team, defining how developers interact with our voice platform as we scale.
$200k - $260k
Staff Platform Engineer, Voice AI
Together AI
Together AI is seeking a Staff Platform Engineer to lead the architecture of their Voice AI platform, which powers real-time voice agents at scale. This role involves setting the technical direction for how developers interact with the platform, from API primitives to autoscaling systems and multi-provider abstractions. The focus is on building robust, low-latency infrastructure for voice applications, which presents unique challenges compared to text inference, such as handling bidirectional audio streams and stateful connections. This is a foundational position on a small team, where decisions will shape the platform's architecture for years to come.
$220k - $280k
Technical Program Manager, Platform
Scale AI
As a Technical Program Manager for the Platform team, you will partner with engineering teams to directly accelerate the development and maturity of the Scale Generative AI Platform (SGP). We are looking for a TPM who has actively built and shipped products in the past and understands how to deliver robust, scalable developer tooling and distributed systems. In this role, you will own the strategic alignment and end-to-end execution of our most critical infrastructure initiatives—from initial scoping to measurable, company-wide and customer-ready adoption. You will serve as the core communication backbone and connective tissue between platform engineering, product teams, and executive leadership. Operating in a hyper-growth, demanding AI environment, you will translate SGP’s architectural complexities into clear execution strategies, unblock engineering bottlenecks, proactively mitigate deployment risks, and ensure our foundational platforms deliver reliable, performant, and secure systems capable of global-scale deployment.
$211k - $264k
Senior Software Engineer, Public Sector
Scale AI
Scale is seeking Senior Software Engineers to join our Public Sector team. You will build core product components that enable forward-deployed teams to develop agentic capabilities across multiple domains. This involves creating systems to ingest and process federal datasets for real-time decision-making in challenging environments. You will lead the development of new agentic capabilities, including multi-layered guardrails, optimized data retrieval, orchestration of asynchronous agents, automated deviation alerts, and decision-path illustration.
$162k - $311k
Member of Technical Staff - Mid-Training Infra
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible. We build open models that empower individuals to control their intelligence and shape the future of AI. As a Member of Technical Staff focused on Mid-Training Infrastructure, you will be instrumental in designing, building, and operating large-scale GPU infrastructure crucial for high-throughput model inference and mid-training workloads. This role involves developing systems that support synthetic data generation and reinforcement learning pipelines at scale, as well as building high-performance inference platforms capable of serving and evaluating models across thousands of GPUs.
Product Manager, Platform Experience & Developer Product
Cohere
Cohere is seeking a Platform Experience and Developer Product Manager to lead the product strategy for how developers and enterprise technical teams build on, integrate with, and operate Cohere's AI model platform. This role spans managed services, APIs, SDKs, and developer tooling, focusing on creating seamless and reliable experiences for users. You will own Cohere's managed service offerings, defining product thinking around deployment models, data residency, and operational controls. Additionally, you will shape the roadmap for APIs and SDKs, ensuring they are stable, well-documented, and easy to use, and oversee developer experience surfaces like the API console, credential management, and observability tools.
Audio Inference Engineer, Model Efficiency
Cohere
Cohere is seeking an Audio Inference Engineer focused on Model Efficiency to join a fast-growing team of researchers and engineers. The mission of this team is to build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques. As an engineer on this team, you will advance core audio model serving metrics, including latency, throughput, and quality by diving deep into systems, identifying bottlenecks, and delivering creative solutions for audio processing and streaming workloads. You will collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real-time and streaming audio inference.
Machine Learning Infrastructure Engineer, Model Inference
Abridge
Abridge is seeking an ML Infrastructure Engineer, Model Inference to build and optimize the core inference infrastructure powering their machine learning models. This role is crucial for enhancing the scalability, efficiency, and performance of Abridge's AI-driven healthcare solutions. The engineer will collaborate with Infrastructure and Research teams to build, deploy, optimize, and orchestrate AI models, working on a platform that transforms patient-clinician conversations into structured clinical notes in real-time.