Model Serving Jobs

74 open roles mentioning Model Serving

Social and Community Manager

2mo ago
f

fireworks ai

Fireworks is seeking a Social and Community Manager to join their team. This role involves staying ahead of the rapidly evolving AI infrastructure landscape, identifying key trends, and shaping narratives. You will be responsible for making data-driven decisions, executing consistently across social and community channels, and transforming complex AI topics into shareable content. This is a hands-on position requiring strong judgment, creativity, and the ability to proactively engage and grow communities.

San Mateo hybrid FullTime
PyTorchHugging FaceModel Serving +1 more

Partnerships Lead

2mo ago
f

fireworks ai

Fireworks AI is seeking a high-ownership sales operator to drive sourced pipeline and revenue through the Microsoft Azure channel. This quota-carrying, field-facing role involves mapping Microsoft's ISV and enterprise field organization, activating co-sell motions, and building repeatable sales plays to guide customers to Fireworks AI through Azure Foundry. You will own the forecast and the number, with significant influence over how Fireworks engages Microsoft's field teams and builds a scalable co-sell motion. If you thrive on creating from scratch and are accountable to a number, this role is for you.

New York hybrid FullTime
AWSAzurePyTorch +4 more

Machine Learning Engineer, Reliability

2mo ago
Fal

Fal

fal is building the generative media ecosystem for the next generation of AI products, providing the infrastructure, tools, and model access needed to scale from idea to production. As generative media reshapes industries, fal is becoming the foundation for ambitious teams. This hybrid ML Engineering / Site Reliability Engineering role will own the reliability, security, and safety of fal's generative media model APIs, ensuring they remain available, performant, secure, and safe for thousands of developers and enterprises. You will address model-specific failure modes, such as degraded output quality, drift, unsafe generations, and abuse patterns, as critical reliability concerns alongside uptime and latency.

Remote - APAC remote FullTime
KubernetesPythonTransformers +1 more

Member of Technical Staff, Cloud Infrastructure, Singapore

2mo ago
f

fireworks ai

Fireworks is building the future of generative AI infrastructure, offering a platform with high-quality models and fast, scalable inference. We are seeking a Backend Software Engineer to design and develop the core backend systems that power our high-performance generative AI platform. Your work will focus on ensuring efficiency, scalability, and stability in handling AI workloads, contributing to cutting-edge AI infrastructure development.

Singapore onsite FullTime
PyTorchModel ServingEmbeddings

ML Engineer, Inference & Optimization

2mo ago
pika

pika

We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale. You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.

Palo Alto HQ onsite FullTime
Deep LearningDistributed TrainingModel Serving

Staff Technical Program Manager

2mo ago
c

crusoe

Crusoe is building the future of AI infrastructure with an energy-first approach, creating a vertically integrated, sustainable AI cloud. We own and operate every layer of the stack, from power generation to AI workloads, to meet the boundless demand for AI compute. As a rapidly growing company, we are seeking problem-solving individuals who are driven by ambition and eager to shape the future of AI infrastructure. The Staff Technical Program Manager will play a crucial role in our Managed Inference platform, a fast-growing product area enabling customers to run production LLM workloads without managing underlying infrastructure. This role involves connecting model engineering, IaaS, product, and data center operations to deliver a reliable and scalable inference platform. You will have a unique opportunity to shape the TPM function as it is still being built, owning end-to-end program delivery, model onboarding, inference optimization, and production readiness for new model versions. Deep familiarity with how LLMs are served, optimized, and evaluated in production is essential.

Tel Aviv - IL onsite FullTime
Fine-TuningModel Serving

Senior Machine Learning Engineer, Voice AI

3mo ago
Together AI

Together AI

Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Senior ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro to achieve frontier-level latency and throughput. You will focus on unique voice inference challenges such as streaming audio, tokenization, and real-time latency budgets, shaping how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.

$200k - $260k

San Francisco remote
PythonGoFine-Tuning +10 more

Staff Platform Engineer, Voice AI

3mo ago
Together AI

Together AI

Together AI is seeking a Staff Platform Engineer to lead the architecture of their Voice AI platform, which powers real-time voice agents at scale. This role involves setting the technical direction for how developers interact with the platform, from API primitives to autoscaling systems and multi-provider abstractions. The focus is on building robust, low-latency infrastructure for voice applications, which presents unique challenges compared to text inference, such as handling bidirectional audio streams and stateful connections. This is a foundational position on a small team, where decisions will shape the platform's architecture for years to come.

$220k - $280k

San Francisco remote
KubernetesPythonTypeScript +7 more

Senior Platform Engineer, Voice AI

3mo ago
Together AI

Together AI

Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications with best-in-class latency and reliability. We are seeking a Senior Platform Engineer to take ownership of the API and infrastructure layer for voice workloads. You will develop the real-time WebSocket and HTTP APIs used by developers to deploy voice experiences, design autoscaling for latency-sensitive streaming workloads, and ensure the reliability of our multi-provider voice platform for production voice agents handling millions of calls. This is a critical, foundational role on a small, high-impact team, defining how developers interact with our voice platform as we scale.

$200k - $260k

San Francisco remote
KubernetesPythonTypeScript +8 more

Staff Machine Learning Engineer, Voice AI

3mo ago
Together AI

Together AI

Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Staff ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro, focusing on pushing latency and throughput boundaries. You will address unique challenges in voice inference, such as streaming audio and real-time latency, and shape the future of how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.

$220k - $280k

San Francisco remote
PythonGoFine-Tuning +10 more

LLM Inference Frameworks and Optimization Engineer

3mo ago
Together AI

Together AI

Together.ai is building state-of-the-art infrastructure for efficient and scalable inference of large language models (LLMs). The company's mission is to optimize inference frameworks, algorithms, and infrastructure to push the boundaries of performance, scalability, and cost-efficiency. They are seeking an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines for multimodal and language models at scale. This role will focus on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design, ensuring efficient large-scale deployment of LLMs and vision models. This position offers a unique opportunity to shape the future of LLM inference infrastructure and ensure scalable, high-performance AI deployment across diverse applications.

$160k - $230k

San Francisco, Singapore, Amsterdam remote
KubernetesPythonC# +9 more

Technical Program Manager, Platform

3mo ago
Scale AI

Scale AI

As a Technical Program Manager for the Platform team, you will partner with engineering teams to directly accelerate the development and maturity of the Scale Generative AI Platform (SGP). We are looking for a TPM who has actively built and shipped products in the past and understands how to deliver robust, scalable developer tooling and distributed systems. In this role, you will own the strategic alignment and end-to-end execution of our most critical infrastructure initiatives—from initial scoping to measurable, company-wide and customer-ready adoption. You will serve as the core communication backbone and connective tissue between platform engineering, product teams, and executive leadership. Operating in a hyper-growth, demanding AI environment, you will translate SGP’s architectural complexities into clear execution strategies, unblock engineering bottlenecks, proactively mitigate deployment risks, and ensure our foundational platforms deliver reliable, performant, and secure systems capable of global-scale deployment.

$211k - $264k

San Francisco, CA; New York, NY onsite
AWSKubernetesGCP +8 more

Senior Software Engineer, Public Sector

3mo ago
Scale AI

Scale AI

Scale is seeking Senior Software Engineers to join our Public Sector team. You will build core product components that enable forward-deployed teams to develop agentic capabilities across multiple domains. This involves creating systems to ingest and process federal datasets for real-time decision-making in challenging environments. You will lead the development of new agentic capabilities, including multi-layered guardrails, optimized data retrieval, orchestration of asynchronous agents, automated deviation alerts, and decision-path illustration.

$162k - $311k

San Francisco, CA; St. Louis, MO; New York, NY; Washington, DC onsite
AWSAzureDocker +11 more

Founding AI Solutions Engineer - USA

3mo ago
i

inworld

Inworld is seeking a Founding AI Solutions Engineer to bridge the gap between sales, product, and engineering. This role is crucial for the revenue team, working directly with senior leadership and the Go-To-Market (GTM) team to guide enterprise and developer customers from initial interest to production deployment. You will be responsible for running Proofs of Concept (POCs), building prototypes, and translating complex technical concepts into clear explanations for diverse audiences. This position offers a unique opportunity to shape Inworld's scaling across various industries and define the Solutions Engineering function as an early hire.

$170k - $250k

Mountain View, California, USA onsite FullTime
PythonTypeScriptJavaScript +2 more

Staff Engineer, API Platform

3mo ago
s

sarvam

Sarvam is building India's full-stack sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. We partner with leading enterprises and public institutions, backed by prominent venture capital firms. Our model APIs, handling millions of daily calls, are crucial for developers and enterprises shipping products on our foundation models. Currently built with FastAPI and Python, we are undertaking a significant rewrite in Go. We are seeking a Staff Engineer to take end-to-end ownership of this critical platform, encompassing its architecture, reliability, performance, and standards.

Bengaluru onsite FullTime
KubernetesPythonGo +4 more

Staff Technical Program Manager, Managed Intelligence

3mo ago
c

crusoe

Crusoe is seeking a Staff Technical Program Manager to join their Managed Intelligence team. This role is crucial for connecting model engineering, IaaS, product, and data center operations to deliver a reliable and scalable inference platform for AI-native companies. You will own end-to-end program delivery, including multi-quarter roadmaps, model onboarding, inference optimization, and production readiness for new model versions. This is a unique opportunity to shape the TPM function within a rapidly growing, vertically integrated AI infrastructure company.

$200k - $240k

San Francisco, CA - US onsite FullTime
Fine-TuningModel Serving

Senior Research Engineer, Audio and Speech

3mo ago
d

decagon

Decagon is seeking a Senior Research Engineer focused on Audio and Speech to build and deploy the next generation of AI voice agents. This role involves advancing multimodal and full-duplex systems that can listen, reason, speak, and respond naturally in real time. You will own your work end-to-end, shipping real improvements and making high-impact technical decisions in a fast-paced, in-office environment.

$200k - $400k

San Francisco onsite FullTime
PythonAI AgentsPyTorch +1 more

Associate Product Manager

4mo ago
f

fireworks ai

This is a rare opportunity for an early-career product manager to work at the frontier of AI infrastructure, building tools used daily by developers and AI teams at the world's most ambitious companies. Inspired by Google’s APM program, APMs at Fireworks will receive mentorship from experienced PMs while also getting rotations across both Fireworks’ product areas and tasks. This role is designed to help you develop the foundational skills of a start-up product leader: rigorous thinking, user empathy, technical depth, and cross-functional leadership.

San Mateo hybrid FullTime
Fine-TuningPyTorchModel Serving +1 more

Software Engineer - Voice AI (Inference Runtime)

4mo ago
B

Baseten

Baseten is seeking a highly impactful individual to lead the Voice AI product area, owning the end-to-end development and implementation of their in-house inference stack for Voice AI models. This role involves partnering closely with various engineering teams to push the boundaries of Voice AI, making a significant impact on industries like productivity, customer service, and education. You will be responsible for bringing state-of-the-art open-source models into production, focusing on optimizing model serving for latency, throughput, and GPU efficiency, and building large-scale, real-time infrastructure for multi-model voice agents.

San Francisco hybrid FullTime
DockerKubernetesPython +5 more

ML Ops Engineer, Chanakya

4mo ago
s

sarvam

Sarvam is building India's full-stack sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. The MLOps Engineer will own the model lifecycle across all deployments, ensuring systems are always operational, accurate, and auditable. This role involves supporting field engineers and managing deployment infrastructure for new products, with an uncompromising standard for reliability, as model failures are considered operational risks.

Delhi onsite FullTime
DockerKubernetesPython +3 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.