Model Serving Jobs
74 open roles mentioning Model Serving
Social and Community Manager
fireworks ai
Fireworks is seeking a Social and Community Manager to join their team. This role involves staying ahead of the rapidly evolving AI infrastructure landscape, identifying key trends, and shaping narratives. You will be responsible for making data-driven decisions, executing consistently across social and community channels, and transforming complex AI topics into shareable content. This is a hands-on position requiring strong judgment, creativity, and the ability to proactively engage and grow communities.
Partnerships Lead
fireworks ai
Fireworks AI is seeking a high-ownership sales operator to drive sourced pipeline and revenue through the Microsoft Azure channel. This quota-carrying, field-facing role involves mapping Microsoft's ISV and enterprise field organization, activating co-sell motions, and building repeatable sales plays to guide customers to Fireworks AI through Azure Foundry. You will own the forecast and the number, with significant influence over how Fireworks engages Microsoft's field teams and builds a scalable co-sell motion. If you thrive on creating from scratch and are accountable to a number, this role is for you.
Machine Learning Engineer, Reliability
Fal
fal is building the generative media ecosystem for the next generation of AI products, providing the infrastructure, tools, and model access needed to scale from idea to production. As generative media reshapes industries, fal is becoming the foundation for ambitious teams. This hybrid ML Engineering / Site Reliability Engineering role will own the reliability, security, and safety of fal's generative media model APIs, ensuring they remain available, performant, secure, and safe for thousands of developers and enterprises. You will address model-specific failure modes, such as degraded output quality, drift, unsafe generations, and abuse patterns, as critical reliability concerns alongside uptime and latency.
Member of Technical Staff, Cloud Infrastructure, Singapore
fireworks ai
Fireworks is building the future of generative AI infrastructure, offering a platform with high-quality models and fast, scalable inference. We are seeking a Backend Software Engineer to design and develop the core backend systems that power our high-performance generative AI platform. Your work will focus on ensuring efficiency, scalability, and stability in handling AI workloads, contributing to cutting-edge AI infrastructure development.
ML Engineer, Inference & Optimization
pika
We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale. You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.
Staff Technical Program Manager
crusoe
Crusoe is building the future of AI infrastructure with an energy-first approach, creating a vertically integrated, sustainable AI cloud. We own and operate every layer of the stack, from power generation to AI workloads, to meet the boundless demand for AI compute. As a rapidly growing company, we are seeking problem-solving individuals who are driven by ambition and eager to shape the future of AI infrastructure. The Staff Technical Program Manager will play a crucial role in our Managed Inference platform, a fast-growing product area enabling customers to run production LLM workloads without managing underlying infrastructure. This role involves connecting model engineering, IaaS, product, and data center operations to deliver a reliable and scalable inference platform. You will have a unique opportunity to shape the TPM function as it is still being built, owning end-to-end program delivery, model onboarding, inference optimization, and production readiness for new model versions. Deep familiarity with how LLMs are served, optimized, and evaluated in production is essential.
Senior Machine Learning Engineer, Voice AI
Together AI
Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Senior ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro to achieve frontier-level latency and throughput. You will focus on unique voice inference challenges such as streaming audio, tokenization, and real-time latency budgets, shaping how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.
$200k - $260k
Staff Platform Engineer, Voice AI
Together AI
Together AI is seeking a Staff Platform Engineer to lead the architecture of their Voice AI platform, which powers real-time voice agents at scale. This role involves setting the technical direction for how developers interact with the platform, from API primitives to autoscaling systems and multi-provider abstractions. The focus is on building robust, low-latency infrastructure for voice applications, which presents unique challenges compared to text inference, such as handling bidirectional audio streams and stateful connections. This is a foundational position on a small team, where decisions will shape the platform's architecture for years to come.
$220k - $280k
Senior Platform Engineer, Voice AI
Together AI
Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications with best-in-class latency and reliability. We are seeking a Senior Platform Engineer to take ownership of the API and infrastructure layer for voice workloads. You will develop the real-time WebSocket and HTTP APIs used by developers to deploy voice experiences, design autoscaling for latency-sensitive streaming workloads, and ensure the reliability of our multi-provider voice platform for production voice agents handling millions of calls. This is a critical, foundational role on a small, high-impact team, defining how developers interact with our voice platform as we scale.
$200k - $260k
Staff Machine Learning Engineer, Voice AI
Together AI
Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications. We are seeking a Staff ML Engineer to lead the model serving layer for voice workloads. This role involves hands-on optimization of inference engines and models like Whisper, Parakeet, Orpheus, and Kokoro, focusing on pushing latency and throughput boundaries. You will address unique challenges in voice inference, such as streaming audio and real-time latency, and shape the future of how voice models are served as the industry shifts towards end-to-end speech-to-speech systems. This is a foundational hire on a small, high-impact team.
$220k - $280k
LLM Inference Frameworks and Optimization Engineer
Together AI
Together.ai is building state-of-the-art infrastructure for efficient and scalable inference of large language models (LLMs). The company's mission is to optimize inference frameworks, algorithms, and infrastructure to push the boundaries of performance, scalability, and cost-efficiency. They are seeking an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines for multimodal and language models at scale. This role will focus on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design, ensuring efficient large-scale deployment of LLMs and vision models. This position offers a unique opportunity to shape the future of LLM inference infrastructure and ensure scalable, high-performance AI deployment across diverse applications.
$160k - $230k
Technical Program Manager, Platform
Scale AI
As a Technical Program Manager for the Platform team, you will partner with engineering teams to directly accelerate the development and maturity of the Scale Generative AI Platform (SGP). We are looking for a TPM who has actively built and shipped products in the past and understands how to deliver robust, scalable developer tooling and distributed systems. In this role, you will own the strategic alignment and end-to-end execution of our most critical infrastructure initiatives—from initial scoping to measurable, company-wide and customer-ready adoption. You will serve as the core communication backbone and connective tissue between platform engineering, product teams, and executive leadership. Operating in a hyper-growth, demanding AI environment, you will translate SGP’s architectural complexities into clear execution strategies, unblock engineering bottlenecks, proactively mitigate deployment risks, and ensure our foundational platforms deliver reliable, performant, and secure systems capable of global-scale deployment.
$211k - $264k
Senior Software Engineer, Public Sector
Scale AI
Scale is seeking Senior Software Engineers to join our Public Sector team. You will build core product components that enable forward-deployed teams to develop agentic capabilities across multiple domains. This involves creating systems to ingest and process federal datasets for real-time decision-making in challenging environments. You will lead the development of new agentic capabilities, including multi-layered guardrails, optimized data retrieval, orchestration of asynchronous agents, automated deviation alerts, and decision-path illustration.
$162k - $311k
Founding AI Solutions Engineer - USA
inworld
Inworld is seeking a Founding AI Solutions Engineer to bridge the gap between sales, product, and engineering. This role is crucial for the revenue team, working directly with senior leadership and the Go-To-Market (GTM) team to guide enterprise and developer customers from initial interest to production deployment. You will be responsible for running Proofs of Concept (POCs), building prototypes, and translating complex technical concepts into clear explanations for diverse audiences. This position offers a unique opportunity to shape Inworld's scaling across various industries and define the Solutions Engineering function as an early hire.
$170k - $250k
Staff Engineer, API Platform
sarvam
Sarvam is building India's full-stack sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. We partner with leading enterprises and public institutions, backed by prominent venture capital firms. Our model APIs, handling millions of daily calls, are crucial for developers and enterprises shipping products on our foundation models. Currently built with FastAPI and Python, we are undertaking a significant rewrite in Go. We are seeking a Staff Engineer to take end-to-end ownership of this critical platform, encompassing its architecture, reliability, performance, and standards.
Staff Technical Program Manager, Managed Intelligence
crusoe
Crusoe is seeking a Staff Technical Program Manager to join their Managed Intelligence team. This role is crucial for connecting model engineering, IaaS, product, and data center operations to deliver a reliable and scalable inference platform for AI-native companies. You will own end-to-end program delivery, including multi-quarter roadmaps, model onboarding, inference optimization, and production readiness for new model versions. This is a unique opportunity to shape the TPM function within a rapidly growing, vertically integrated AI infrastructure company.
$200k - $240k
Senior Research Engineer, Audio and Speech
decagon
Decagon is seeking a Senior Research Engineer focused on Audio and Speech to build and deploy the next generation of AI voice agents. This role involves advancing multimodal and full-duplex systems that can listen, reason, speak, and respond naturally in real time. You will own your work end-to-end, shipping real improvements and making high-impact technical decisions in a fast-paced, in-office environment.
$200k - $400k
Associate Product Manager
fireworks ai
This is a rare opportunity for an early-career product manager to work at the frontier of AI infrastructure, building tools used daily by developers and AI teams at the world's most ambitious companies. Inspired by Google’s APM program, APMs at Fireworks will receive mentorship from experienced PMs while also getting rotations across both Fireworks’ product areas and tasks. This role is designed to help you develop the foundational skills of a start-up product leader: rigorous thinking, user empathy, technical depth, and cross-functional leadership.
Software Engineer - Voice AI (Inference Runtime)
Baseten
Baseten is seeking a highly impactful individual to lead the Voice AI product area, owning the end-to-end development and implementation of their in-house inference stack for Voice AI models. This role involves partnering closely with various engineering teams to push the boundaries of Voice AI, making a significant impact on industries like productivity, customer service, and education. You will be responsible for bringing state-of-the-art open-source models into production, focusing on optimizing model serving for latency, throughput, and GPU efficiency, and building large-scale, real-time infrastructure for multi-model voice agents.
ML Ops Engineer, Chanakya
sarvam
Sarvam is building India's full-stack sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. The MLOps Engineer will own the model lifecycle across all deployments, ensuring systems are always operational, accurate, and auditable. This role involves supporting field engineers and managing deployment infrastructure for new products, with an uncompromising standard for reliability, as model failures are considered operational risks.