Model Serving Jobs

74 open roles mentioning Model Serving

Staff Software Engineer, AI Reliability

24d ago
Anthropic

Anthropic

Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. The AI Reliability Engineering (AIRE) team partners with other teams across Anthropic to enhance the reliability of critical serving paths, from SDKs through API layers, serving infrastructure, and accelerators. This role offers a unique, cross-cutting exposure to the most important systems at Anthropic, requiring a holistic view of system composition and reliability.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicClaudeModel Serving

Product Manager, Platform

25d ago
f

fireworks ai

Fireworks is seeking a Product Manager to focus on their core inference product and general platform. This role is central to enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. You will be responsible for setting strategy, writing product specifications, engaging directly with customers running production traffic, and collaborating across inference, infrastructure, and go-to-market teams to deliver significant customer impact. Example challenges include optimizing rate limits and preventing fraudulent usage while maintaining a smooth experience for legitimate users.

San Mateo hybrid FullTime
GoPyTorchModel Serving +1 more

Product Manager, Training

25d ago
f

fireworks ai

Fireworks is seeking a Product Manager to focus on their AI training platform. In this role, you will be instrumental in shaping the strategy, defining product specifications, and directly engaging with customers to understand their real-world tuning challenges. You will collaborate closely with the training, research, and inference teams to deliver impactful solutions that enhance customer AI development. Example initiatives include driving growth in self-serve training usage and improving the speed and ease of training for enterprise clients.

San Mateo hybrid FullTime
Fine-TuningPyTorchModel Serving +2 more

Staff Engineer - Mobile, Desktop & KMP

26d ago
s

sarvam

Sarvam is building India's sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. The company partners with leading enterprises and public institutions. This role focuses on Sarvam's consumer and enterprise products, which are native applications running on devices like Android phones, iPhones, Windows laptops, Linux desktops, and Macs. These applications require on-device inference, audio capture/streaming, and offline functionality, needing to perform reliably across a wide range of hardware and network conditions in India. The Staff Engineer will own the client application layer across mobile and desktop, using Kotlin Multiplatform (KMP) as a unifying technology to manage development across five platforms without separate codebases. This is an individual contributor role where you will collaborate closely with backend, model serving, and product teams.

Bengaluru onsite FullTime
GoTensorFlowModel Serving

Research Engineer, LangSmith Engine

1mo ago
L

Langchain

LangChain is seeking an experienced Research Engineer to join the LangSmith Engine team. This role focuses on enhancing the capabilities and efficiency of a proactive agent engineer that analyzes production traces, identifies failures, and implements fixes to prevent recurrence. You will study agent failures, build benchmarks, run experiments to improve performance, and translate successful ideas into production. This involves optimizing prompting, agent harnesses, model selection, fine-tuning, and post-training custom models, with a strong emphasis on measurable improvements to the overall agent. The role also requires an understanding of production engineering and system-level trade-offs, including cost, latency, reliability, and scalability, working closely with production engineers to ensure reliable real-world performance.

New York, NY onsite FullTime
LangGraphLangChainAI Agents +4 more

Partnerships Lead

1mo ago
f

fireworks ai

Fireworks AI is seeking a high-ownership sales operator to drive sourced pipeline and revenue through the Microsoft Azure channel. This quota-carrying, field-facing role involves mapping Microsoft's ISV and enterprise field organization, activating co-sell motions with key personnel, and building repeatable sales plays to guide customers to Fireworks AI via Azure Foundry. You will own the forecast and the number, with significant influence over how Fireworks engages Microsoft's field teams and builds a scalable co-sell motion. This is an opportunity for someone who thrives on creating from scratch and is accountable to achieving targets in a dynamic environment.

London hybrid FullTime
AWSAzurePyTorch +4 more

Product Designer

1mo ago
f

fireworks ai

Fireworks is seeking a Product Designer passionate about transforming complex AI infrastructure into intuitive developer experiences. This role involves owning the design process end-to-end, from model deployment to AI agent creation on the Fireworks platform. The focus is on simplifying powerful technologies, requiring close collaboration with engineers and PMs who have a deep understanding of systems. The ideal candidate will be excited about console UIs, API ergonomics, and enhancing the developer's journey to first inference.

San Mateo hybrid FullTime
AWSAzureAI Agents +4 more

Engineering Site Lead

1mo ago
P

Perplexity AI

Perplexity is seeking an exceptional Site Lead to establish and scale its London office, a strategic presence in one of the world's leading tech hubs. This role involves building teams and culture from the ground up while driving technical excellence in infrastructure and AI systems. The Site Lead will serve as the face of Perplexity in London, responsible for building the technical organization, fostering a world-class engineering culture, and directly managing infrastructure teams. This position reports to senior leadership and collaborates cross-functionally with global teams.

$15k - $30k

London onsite FullTime
AWSAzureKubernetes +3 more

Performance Engineer, Inference

1mo ago
s

sarvam

Sarvam is seeking a Performance Engineer specializing in Inference to join their Performance Engineering team. This role focuses on integrating and optimizing model serving stacks for large, distributed models across a fleet of GPUs. You will be responsible for the end-to-end production serving path, modifying and extending existing serving runtimes like SGLang, vLLM, NVIDIA Dynamo, or TensorRT-LLM. The position involves building and training custom speculative decoding models and ensuring the performance and cost-efficiency of the serving infrastructure. You will collaborate closely with model, kernel, and SRE teams to achieve critical performance metrics such as latency, throughput, and GPU utilization.

Bengaluru onsite FullTime
C#Model Serving

GTM Engineer

1mo ago
f

fireworks ai

Fireworks is seeking a GTM Engineer to design, operate, and enhance the systems that drive our revenue engine. This role bridges GTM systems architecture and field execution, ensuring our tools and automations boost productivity for sales and marketing teams. You will collaborate closely with sellers, marketers, and GTM leadership to ensure our systems provide crucial signals, reduce manual tasks, and accelerate team velocity. The ideal candidate is passionate about inventing the future of sales platforms and applying world-class AI to go-to-market execution.

San Francisco hybrid FullTime
GoPyTorchModel Serving +1 more

Member of Technical Staff, Research

1mo ago
f

fireworks ai

Fireworks is seeking a Member of Technical Staff for its Research team to push the boundaries of generative AI. This role involves advancing LLMs and multimodal systems through foundational research, focusing on enhancing model efficiency, accuracy, and scalability to shape high-performance AI infrastructure. You will collaborate with experts in deep learning, distributed systems, and optimization to translate cutting-edge research into practical applications and influence how leading companies build and deploy AI.

San Mateo hybrid FullTime
PythonC#PyTorch +5 more

Member of Technical Staff, AI Training Infrastructure

1mo ago
f

fireworks ai

Fireworks is seeking a Training Infrastructure Engineer to design, build, and optimize the infrastructure that powers large-scale model training operations. This role is crucial for developing high-performance AI training infrastructure, requiring collaboration with AI researchers and engineers to create robust training pipelines, optimize distributed training workloads, and ensure reliable model development. The position offers the opportunity to solve hard problems at the forefront of AI infrastructure, build what's next with bleeding-edge technology, and have a direct impact on the future of AI within a fast-growing, passionate team.

San Mateo hybrid FullTime
AWSAzureDocker +5 more

ML Platform Engineer

1mo ago
s

synthesia.io

Synthesia is seeking an Engineer to join its ML Platform team. This team is responsible for building and operating the systems that enable researchers and product teams to train, serve, and deploy generative models efficiently and reliably. The role involves working on research infrastructure, production serving systems, internal tooling, and platform interfaces, with a growing focus on making these systems automation-friendly and agent-oriented. This is a hands-on individual contributor role with significant ownership, where you will help shape the evolution of the ML platform as it scales.

Europe remote FullTime
KubernetesPythonTerraform +1 more

Tech Lead Manager, Inference

1mo ago
l

lumalabs

Luma is seeking a Tech Lead Manager for its Inference team to own the entire inference serving stack, encompassing routing, scheduling, and fleet-wide orchestration across thousands of GPUs, multiple clouds, and hardware vendors. This is a hands-on role where at least half of your time will be dedicated to architecting and building core platform components, making critical design decisions, and debugging complex incidents. You will also be responsible for leading, growing, and developing the inference engineering team, including hiring, coaching, and managing on-call rotations. The role involves setting the technical roadmap for serving infrastructure, owning platform SLOs and economics, and partnering with research to deploy new architectures and integrate serving into online RL and evaluation loops. The ideal candidate has extensive experience operating large-scale inference fleets and a genuine desire to remain hands-on in building and improving the serving stack.

$30k - $60k

Redwood City, CA hybrid FullTime
KubernetesPythonRust +3 more

Software Engineer, Inference

1mo ago
l

lumalabs

Luma is seeking a Software Engineer to own the serving of their models. This role involves integrating new architectures into the inference engine, scaling deployments across thousands of machines, and optimizing GPU fleet utilization while meeting internal service level objectives (SLOs). The work focuses on large-scale inference systems, including scheduling, fleet management, deployment pipelines, and reliability across various clusters and hardware providers. This position is ideal for a strong systems engineer experienced with model serving and Kubernetes at scale, rather than pure modeling.

$30k - $60k

Redwood City, CA hybrid FullTime
DockerKubernetesPython +4 more

Research Scientist / Engineer – Reinforcement Learning Infrastructure

1mo ago
l

lumalabs

Luma is seeking a Research Scientist / Engineer to build the systems that enable reinforcement learning (RL) at frontier scale. This role involves coupling policy optimization with large fleets of inference workers, agentic environments, and reward/verification systems to transform model behavior into learning signals. RL is crucial for Luma's models to evolve from capable to useful. Operating RL at scale is a complex systems challenge, encompassing training, rollout generation, environment execution, and reward computation across thousands of GPUs, demanding speed, stability, and correctness. This position is ideal for someone with hands-on experience in post-training LLMs with RL, building environments and verifiers, and debugging large-scale asynchronous rollout pipelines.

$30k - $60k

Redwood City, CA hybrid FullTime
KubernetesGoPyTorch +3 more

Applied Machine Learning Engineer, Singapore

1mo ago
f

fireworks ai

As an Applied Machine Learning Engineer, you will serve as a vital bridge between cutting-edge AI research and practical, real-world applications. Your work will focus on developing, fine-tuning, and operationalizing machine learning models that drive business value and enhance user experiences. This is a hands-on engineering role that combines deep technical expertise with a strong customer focus to deliver scalable AI solutions.

Singapore onsite FullTime
PythonFine-TuningPyTorch +4 more

Head of GTM Engineering & Systems

2mo ago
f

fireworks ai

Fireworks is seeking a leader for its growing Go-To-Market (GTM) Engineering team. This role will be responsible for defining the architecture and driving the AI roadmap to empower customer-facing teams. You will own the entire GTM technology stack, including Salesforce, CPQ, enrichment, and engagement tools, ensuring seamless integration and scalability for a rapidly expanding organization. A key focus will be bringing agentic AI from concept to production, establishing it as the new operational standard for GTM teams, rather than just an experiment. This is a hands-on builder and leader position, requiring someone who can make critical architectural decisions, deliver quickly, and elevate the performance of the GTM Engineering team.

$50k - $300k

San Mateo hybrid FullTime
PyTorchModel ServingEmbeddings

Senior Customer Success Manager, Managed Inference

2mo ago
c

crusoe

We are seeking a highly motivated and skilled Senior Customer Success Manager with a strong background in customer engagement and a deep technical understanding of cloud computing, AI, and ML. The ideal candidate has experience supporting customers running production AI inference workloads and understands the operational, technical, and business challenges associated with deploying and scaling AI applications. Experience supporting Managed Inference, model serving platforms, LLM deployments, AI agents, or GPU-based inference environments is highly preferred. This role is pivotal in ensuring that our clients maximize the value of our solutions, guiding them through the technical complexities and empowering them with the tools and knowledge to achieve their business and sustainability goals.

$190k - $215k

Denver, CO - US onsite FullTime
KubernetesRAGAI Agents +1 more

Senior Customer Success Manager, Managed Inference

2mo ago
c

crusoe

We are seeking a highly motivated and skilled Senior Customer Success Manager with a strong background in customer engagement and a deep technical understanding of cloud computing, AI, and ML. The ideal candidate has experience supporting customers running production AI inference workloads and understands the operational, technical, and business challenges associated with deploying and scaling AI applications. Experience supporting Managed Inference, model serving platforms, LLM deployments, AI agents, or GPU-based inference environments is highly preferred. This role is pivotal in ensuring that our clients maximize the value of our solutions, guiding them through the technical complexities and empowering them with the tools and knowledge to achieve their business and sustainability goals.

$190k - $215k

San Francisco, CA - US onsite FullTime
KubernetesRAGAI Agents +1 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.