Model Serving Jobs
74 open roles mentioning Model Serving
Member of Technical Staff, Systems Infrastructure (2026 PhD New Grad)
fireworks ai
Fireworks is seeking PhD graduates to join their Systems Infrastructure team. This role is designed for individuals finishing their PhD in Computer Science, Computer Engineering, Electrical Engineering, or a similar field, who are interested in applying their research to large-scale, real-world AI infrastructure. You will be responsible for designing and building the core systems that power Fireworks, including schedulers, storage systems, and networks, to ensure efficient operation of tens of thousands of accelerators and low inference latency. You will tackle complex problems such as optimizing job placement on heterogeneous hardware, high-speed data movement for model weights, and maintaining saturated datacenter networks. You will be paired with a senior engineer mentor and work on impactful projects from day one, with start dates flexible around thesis defense.
Member of Technical Staff, Research (2026 PhD New Grad)
fireworks ai
Fireworks is seeking a Member of Technical Staff on the Research team, designed for PhD candidates who want their work to reach production. This role involves pushing the boundaries of generative AI, advancing LLMs and multimodal systems through foundational research. You will enhance model efficiency, accuracy, and scalability, directly shaping high-performance AI infrastructure. You will be paired with a senior researcher as a mentor and given a real research problem from day one, collaborating with top experts in deep learning, distributed systems, and optimization. The work you contribute will be deployed by leading companies, often within weeks.
Software Engineer - Training/Inference (C++)
xAI
SpaceXAI is building a high-performance inference platform that serves Grok to millions of users daily with exceptional speed and reliability. As a Member of Technical Staff - Inference, you will be responsible for designing and optimizing large-scale model serving systems from end-to-end. This includes distributed infrastructure components like global KV cache, continuous batching, and load balancing, as well as deep low-level optimizations such as GPU kernels, quantization, and speculative decoding. This is a high-impact role where your contributions will directly influence the speed and reliability of user interactions with Grok at a massive scale.
$180k - $440k
Backend Engineer - API
xAI
SpaceXAI is building AI systems to understand the universe and aid humanity. The team is small, highly motivated, and focused on engineering excellence, operating with a flat structure where all employees are hands-on and contribute directly to the mission. This role requires strong communication skills and the ability to share knowledge concisely. The ideal candidate will have a good understanding of building highly scalable and reliable production infrastructure, with familiarity in compiled languages like Rust, C++, or Go being highly beneficial.
£107k - £262k
Member of Technical Staff (EMEA)
fireworks ai
Fireworks is seeking a Member of Technical Staff for their EMEA field team. This role involves building on top of one of the world's largest inference platforms, spanning capabilities on inference and serving, tooling for fine-tuning and post-training of LLMs, and developing product features specific to the region. The ideal candidate is an excellent software engineer who prefers building product, with a strong emphasis on first principles thinking and the ability to reason about unfamiliar problems from the ground up. This is a flexible role designed for exceptional builders with 3 to 10 years of experience.
Compensation Partner
fireworks ai
Fireworks is seeking its first dedicated Compensation Partner to translate the company's compensation philosophy into practical application. This role will build the essential systems and processes for benchmarking, reviewing, and adjusting compensation as the organization grows rapidly. The Compensation Partner will be responsible for hands-on execution, including market research, managing review cycles, ensuring data accuracy, and serving as the primary resource for compensation-related inquiries from HRBPs and leaders. Additionally, this role will collaborate closely with Finance to develop and refine the company's equity framework, bringing a compensation perspective to financial planning.
Security Engineer, Application Security
Mercor
Mercor is a leading AI data company building the layer between human expertise and frontier models. We are seeking a Security Engineer, Application Security to own application security at a company where the app layer is the highest-priority security surface. This is not a scan-and-triage role. You will embed in the development lifecycle, review code for exploitable flaws, build security tooling into CI/CD, and drive vulnerability remediation across a platform serving 300K+ experts and enterprise clients processing sensitive AI training data. You should be comfortable building alongside AI code-gen tools, using LLMs to accelerate code review and threat modeling, and automating away repetitive work.
Research Engineer, Audio and Speech
decagon
Decagon is seeking a Research Engineer focused on Audio and Speech to build and deploy the next generation of AI voice agents. This role involves developing models and agent harnesses for real-time, full-duplex conversational systems that can listen, reason, speak, and respond naturally. You will own projects end-to-end, from idea to production, making high-impact technical decisions and shipping real improvements in AI voice technology.
$200k - $400k
Senior Machine Learning Engineer
Cloudflare
You will help define how machine learning models run across Cloudflare’s global network, from frontier open LLMs and real-time voice models to customer-deployed models served on heterogeneous GPUs and next-generation accelerators. You’ll work with systems engineers, product teams, hardware partners, and AI/ML engineers to bring models into production with low latency, strong reliability, and efficient resource use. This role combines applied ML, inference optimization, evaluation, and production engineering, with a focus on benchmarking models, improving serving performance, validating quality, and building tooling that helps Cloudflare and its customers ship AI applications at Internet scale.
Head of GTM, AI Inference
Cloudflare
Cloudflare is developing a unique AI inference platform, aiming to be more than just a GPU provider. This role will own the go-to-market strategy for this business, acting as a liaison between AI infrastructure, product teams, and the market. The individual will engage with prospects, collaborate with product leadership on positioning and packaging, and develop analytical frameworks to identify target customers and partners. This is a strategic role where you will define and refine the go-to-market playbook, blending financial modeling of inference unit economics with strategic sales conversations.
Senior Manager, Indirect Tax
fireworks ai
Fireworks is seeking an experienced Senior Manager, Indirect Tax to lead and manage the organization’s indirect tax function. This role will be responsible for indirect tax compliance, reporting, planning, audits, and risk management, while serving as a key advisor to Finance, Accounting, Legal, Operations, and other business stakeholders. The successful candidate will bring strong technical expertise in indirect taxation, excellent judgment, and the ability to translate complex tax requirements into practical business solutions. This role is well suited to a tax professional who can operate strategically while maintaining strong oversight of day-to-day compliance and execution.
Applied Machine Learning Engineer, EMEA
fireworks ai
Fireworks is seeking an Applied Machine Learning Engineer for the EMEA region. In this role, you will be the technical owner of customer engagements, embedding within client teams to understand their specific needs and challenges. You will be responsible for the entire lifecycle of a customer's deployment on the Fireworks platform, from initial scoping and model selection to ensuring production readiness, performance, and cost-efficiency. This position emphasizes first principles thinking and requires a deep understanding of software engineering, machine learning techniques, and infrastructure optimization.
Software Engineer, AI/ML Infrastructure
Glean
Glean is seeking software engineers to contribute to building a leading search and assistant product for the workplace. Engineers will work across the technology stack, focusing on areas such as generative AI, RAG, query understanding, document understanding, domain-adapted language models, natural language question-answering, evaluation, and experimentation. This role involves close interaction with customers to understand their challenges and applying the most effective tools to solve them.
$175k - $270k
Product Manager, Managed North
Cohere
Cohere is seeking a Platform Experience and Developer Product Manager to lead the product strategy for how developers and enterprise technical teams build on, integrate with, and operate Cohere's AI model platform. This role spans managed services, APIs, SDKs, and developer tooling, focusing on creating seamless and reliable experiences for users. You will own Cohere's managed service offerings, defining product thinking around deployment models, data residency, and operational controls. Additionally, you will shape the roadmap for APIs and SDKs, ensuring they are stable, well-documented, and easy to use, and oversee developer experience surfaces like the API console, credential management, and observability tools.
Staff Software Engineer, AI Reliability Engineering
Anthropic
AI Reliability Engineering (AIRE) partners with teams across Anthropic to improve reliability across our most critical serving paths, from SDKs through our network, API layers, serving infrastructure, and accelerators. This role involves jumping into the trenches alongside partner teams to make the systems that deliver Claude more robust and resilient, whether during an incident or collaborating on projects. Reliability is viewed as an emergent phenomenon that transcends single team boundaries, requiring a holistic perspective across the entire system.
Staff Software Engineer, AI Reliability Engineering
Anthropic
AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths, from the SDK through our network, API layers, serving infrastructure, and accelerators. This role involves jumping into the trenches alongside partner teams to make the systems that deliver Claude more robust and resilient, be it during an incident or collaborating on projects. Reliability here is an emergent phenomenon that transcends any single team's boundaries, requiring someone to zoom out and look at the whole picture, offering dynamic, cross-cutting exposure to the systems that matter most.
Developer Advocate
fireworks ai
Fireworks is seeking a Developer Advocate to be the voice of their specialized intelligence platform for developers. In this role, you will be responsible for making Fireworks the go-to platform for building, fine-tuning, and serving frontier open models in production. This position requires a dual focus: you must be a hands-on engineer who ships code and understands technical nuances, and a public technical communicator who can clearly articulate complex topics through writing, speaking, and community engagement. You will be at the forefront of a rapidly evolving field, consistently staying ahead of new model releases and making it look repeatable rather than heroic. As one of the first advocates on the team, you will help set the standards for content, launch responses, and team presence.
Customer Success Manager, Managed Inference
crusoe
Crusoe is seeking a motivated Customer Success Manager to support customers running AI inference workloads on their cloud platform. This role, within the Customer Experience organization, will help clients navigate the operational and technical challenges of deploying and scaling AI applications, from model serving and GPU utilization to production readiness. You will foster strong customer relationships, act as a liaison to technical teams, and assist clients in maximizing the value derived from Crusoe's AI and ML solutions. This is a full-time position based in the Bay Area, CA.
$175k - $200k
Member of Technical Staff (Software Engineer, Model Platform)
Perplexity AI
Perplexity AI is seeking a deeply technical software engineer to own and evolve its model serving platform. This mission-critical system connects products and research systems with inference layers, providing a fast and reliable interface across providers. The role involves working at the intersection of distributed systems, AI products, and infrastructure, making it easier for teams to adopt new models, run experiments, and ship dependable products. The ideal candidate will possess strong technical judgment, curiosity about frontier models, and experience designing core abstractions and leading technical decisions.
Staff+ Software Engineer, Claude Managed Agents
Anthropic
Anthropic is seeking experienced backend and distributed systems engineers to join the Agentic Systems team. This team builds Claude Managed Agents, a hosted platform for creating, running, and scaling production agents on Claude. The platform provides essential functionalities like agent loops, sandboxed execution, state management, credential handling, and error recovery, exposed through stable APIs. This role involves driving new features from ideation to general availability, owning systems end-to-end, and collaborating with product, research, developer experience, and go-to-market teams. The ideal candidate will be comfortable tackling complex distributed systems challenges, have a strong product sense for API design, and be motivated by transforming ambiguous ideas into high-quality platform capabilities.