Model Serving Jobs
74 open roles mentioning Model Serving
Staff Software Engineer, AI Reliability
Anthropic
Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. The AI Reliability Engineering (AIRE) team partners with other teams across Anthropic to enhance the reliability of critical serving paths, from SDKs through API layers, serving infrastructure, and accelerators. This role offers a unique, cross-cutting exposure to the most important systems at Anthropic, requiring a holistic view of system composition and reliability.
Product Manager, Platform
fireworks ai
Fireworks is seeking a Product Manager to focus on their core inference product and general platform. This role is central to enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. You will be responsible for setting strategy, writing product specifications, engaging directly with customers running production traffic, and collaborating across inference, infrastructure, and go-to-market teams to deliver significant customer impact. Example challenges include optimizing rate limits and preventing fraudulent usage while maintaining a smooth experience for legitimate users.
Product Manager, Training
fireworks ai
Fireworks is seeking a Product Manager to focus on their AI training platform. In this role, you will be instrumental in shaping the strategy, defining product specifications, and directly engaging with customers to understand their real-world tuning challenges. You will collaborate closely with the training, research, and inference teams to deliver impactful solutions that enhance customer AI development. Example initiatives include driving growth in self-serve training usage and improving the speed and ease of training for enterprise clients.
Staff Engineer - Mobile, Desktop & KMP
sarvam
Sarvam is building India's sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. The company partners with leading enterprises and public institutions. This role focuses on Sarvam's consumer and enterprise products, which are native applications running on devices like Android phones, iPhones, Windows laptops, Linux desktops, and Macs. These applications require on-device inference, audio capture/streaming, and offline functionality, needing to perform reliably across a wide range of hardware and network conditions in India. The Staff Engineer will own the client application layer across mobile and desktop, using Kotlin Multiplatform (KMP) as a unifying technology to manage development across five platforms without separate codebases. This is an individual contributor role where you will collaborate closely with backend, model serving, and product teams.
Research Engineer, LangSmith Engine
Langchain
LangChain is seeking an experienced Research Engineer to join the LangSmith Engine team. This role focuses on enhancing the capabilities and efficiency of a proactive agent engineer that analyzes production traces, identifies failures, and implements fixes to prevent recurrence. You will study agent failures, build benchmarks, run experiments to improve performance, and translate successful ideas into production. This involves optimizing prompting, agent harnesses, model selection, fine-tuning, and post-training custom models, with a strong emphasis on measurable improvements to the overall agent. The role also requires an understanding of production engineering and system-level trade-offs, including cost, latency, reliability, and scalability, working closely with production engineers to ensure reliable real-world performance.
Partnerships Lead
fireworks ai
Fireworks AI is seeking a high-ownership sales operator to drive sourced pipeline and revenue through the Microsoft Azure channel. This quota-carrying, field-facing role involves mapping Microsoft's ISV and enterprise field organization, activating co-sell motions with key personnel, and building repeatable sales plays to guide customers to Fireworks AI via Azure Foundry. You will own the forecast and the number, with significant influence over how Fireworks engages Microsoft's field teams and builds a scalable co-sell motion. This is an opportunity for someone who thrives on creating from scratch and is accountable to achieving targets in a dynamic environment.
Product Designer
fireworks ai
Fireworks is seeking a Product Designer passionate about transforming complex AI infrastructure into intuitive developer experiences. This role involves owning the design process end-to-end, from model deployment to AI agent creation on the Fireworks platform. The focus is on simplifying powerful technologies, requiring close collaboration with engineers and PMs who have a deep understanding of systems. The ideal candidate will be excited about console UIs, API ergonomics, and enhancing the developer's journey to first inference.
Engineering Site Lead
Perplexity AI
Perplexity is seeking an exceptional Site Lead to establish and scale its London office, a strategic presence in one of the world's leading tech hubs. This role involves building teams and culture from the ground up while driving technical excellence in infrastructure and AI systems. The Site Lead will serve as the face of Perplexity in London, responsible for building the technical organization, fostering a world-class engineering culture, and directly managing infrastructure teams. This position reports to senior leadership and collaborates cross-functionally with global teams.
$15k - $30k
Performance Engineer, Inference
sarvam
Sarvam is seeking a Performance Engineer specializing in Inference to join their Performance Engineering team. This role focuses on integrating and optimizing model serving stacks for large, distributed models across a fleet of GPUs. You will be responsible for the end-to-end production serving path, modifying and extending existing serving runtimes like SGLang, vLLM, NVIDIA Dynamo, or TensorRT-LLM. The position involves building and training custom speculative decoding models and ensuring the performance and cost-efficiency of the serving infrastructure. You will collaborate closely with model, kernel, and SRE teams to achieve critical performance metrics such as latency, throughput, and GPU utilization.
GTM Engineer
fireworks ai
Fireworks is seeking a GTM Engineer to design, operate, and enhance the systems that drive our revenue engine. This role bridges GTM systems architecture and field execution, ensuring our tools and automations boost productivity for sales and marketing teams. You will collaborate closely with sellers, marketers, and GTM leadership to ensure our systems provide crucial signals, reduce manual tasks, and accelerate team velocity. The ideal candidate is passionate about inventing the future of sales platforms and applying world-class AI to go-to-market execution.
Member of Technical Staff, Research
fireworks ai
Fireworks is seeking a Member of Technical Staff for its Research team to push the boundaries of generative AI. This role involves advancing LLMs and multimodal systems through foundational research, focusing on enhancing model efficiency, accuracy, and scalability to shape high-performance AI infrastructure. You will collaborate with experts in deep learning, distributed systems, and optimization to translate cutting-edge research into practical applications and influence how leading companies build and deploy AI.
Member of Technical Staff, AI Training Infrastructure
fireworks ai
Fireworks is seeking a Training Infrastructure Engineer to design, build, and optimize the infrastructure that powers large-scale model training operations. This role is crucial for developing high-performance AI training infrastructure, requiring collaboration with AI researchers and engineers to create robust training pipelines, optimize distributed training workloads, and ensure reliable model development. The position offers the opportunity to solve hard problems at the forefront of AI infrastructure, build what's next with bleeding-edge technology, and have a direct impact on the future of AI within a fast-growing, passionate team.
ML Platform Engineer
synthesia.io
Synthesia is seeking an Engineer to join its ML Platform team. This team is responsible for building and operating the systems that enable researchers and product teams to train, serve, and deploy generative models efficiently and reliably. The role involves working on research infrastructure, production serving systems, internal tooling, and platform interfaces, with a growing focus on making these systems automation-friendly and agent-oriented. This is a hands-on individual contributor role with significant ownership, where you will help shape the evolution of the ML platform as it scales.
Tech Lead Manager, Inference
lumalabs
Luma is seeking a Tech Lead Manager for its Inference team to own the entire inference serving stack, encompassing routing, scheduling, and fleet-wide orchestration across thousands of GPUs, multiple clouds, and hardware vendors. This is a hands-on role where at least half of your time will be dedicated to architecting and building core platform components, making critical design decisions, and debugging complex incidents. You will also be responsible for leading, growing, and developing the inference engineering team, including hiring, coaching, and managing on-call rotations. The role involves setting the technical roadmap for serving infrastructure, owning platform SLOs and economics, and partnering with research to deploy new architectures and integrate serving into online RL and evaluation loops. The ideal candidate has extensive experience operating large-scale inference fleets and a genuine desire to remain hands-on in building and improving the serving stack.
$30k - $60k
Software Engineer, Inference
lumalabs
Luma is seeking a Software Engineer to own the serving of their models. This role involves integrating new architectures into the inference engine, scaling deployments across thousands of machines, and optimizing GPU fleet utilization while meeting internal service level objectives (SLOs). The work focuses on large-scale inference systems, including scheduling, fleet management, deployment pipelines, and reliability across various clusters and hardware providers. This position is ideal for a strong systems engineer experienced with model serving and Kubernetes at scale, rather than pure modeling.
$30k - $60k
Research Scientist / Engineer – Reinforcement Learning Infrastructure
lumalabs
Luma is seeking a Research Scientist / Engineer to build the systems that enable reinforcement learning (RL) at frontier scale. This role involves coupling policy optimization with large fleets of inference workers, agentic environments, and reward/verification systems to transform model behavior into learning signals. RL is crucial for Luma's models to evolve from capable to useful. Operating RL at scale is a complex systems challenge, encompassing training, rollout generation, environment execution, and reward computation across thousands of GPUs, demanding speed, stability, and correctness. This position is ideal for someone with hands-on experience in post-training LLMs with RL, building environments and verifiers, and debugging large-scale asynchronous rollout pipelines.
$30k - $60k
Applied Machine Learning Engineer, Singapore
fireworks ai
As an Applied Machine Learning Engineer, you will serve as a vital bridge between cutting-edge AI research and practical, real-world applications. Your work will focus on developing, fine-tuning, and operationalizing machine learning models that drive business value and enhance user experiences. This is a hands-on engineering role that combines deep technical expertise with a strong customer focus to deliver scalable AI solutions.
Head of GTM Engineering & Systems
fireworks ai
Fireworks is seeking a leader for its growing Go-To-Market (GTM) Engineering team. This role will be responsible for defining the architecture and driving the AI roadmap to empower customer-facing teams. You will own the entire GTM technology stack, including Salesforce, CPQ, enrichment, and engagement tools, ensuring seamless integration and scalability for a rapidly expanding organization. A key focus will be bringing agentic AI from concept to production, establishing it as the new operational standard for GTM teams, rather than just an experiment. This is a hands-on builder and leader position, requiring someone who can make critical architectural decisions, deliver quickly, and elevate the performance of the GTM Engineering team.
$50k - $300k
Senior Customer Success Manager, Managed Inference
crusoe
We are seeking a highly motivated and skilled Senior Customer Success Manager with a strong background in customer engagement and a deep technical understanding of cloud computing, AI, and ML. The ideal candidate has experience supporting customers running production AI inference workloads and understands the operational, technical, and business challenges associated with deploying and scaling AI applications. Experience supporting Managed Inference, model serving platforms, LLM deployments, AI agents, or GPU-based inference environments is highly preferred. This role is pivotal in ensuring that our clients maximize the value of our solutions, guiding them through the technical complexities and empowering them with the tools and knowledge to achieve their business and sustainability goals.
$190k - $215k
Senior Customer Success Manager, Managed Inference
crusoe
We are seeking a highly motivated and skilled Senior Customer Success Manager with a strong background in customer engagement and a deep technical understanding of cloud computing, AI, and ML. The ideal candidate has experience supporting customers running production AI inference workloads and understands the operational, technical, and business challenges associated with deploying and scaling AI applications. Experience supporting Managed Inference, model serving platforms, LLM deployments, AI agents, or GPU-based inference environments is highly preferred. This role is pivotal in ensuring that our clients maximize the value of our solutions, guiding them through the technical complexities and empowering them with the tools and knowledge to achieve their business and sustainability goals.
$190k - $215k