Model Serving Jobs

74 open roles mentioning Model Serving

Strategic Deployment Engineer, Chanakya

4mo ago
s

sarvam

Sarvam is building India's full-stack sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. As a Strategic Deployment Engineer, you will be embedded with clients, owning the entire lifecycle of AI system deployments in complex and often constrained environments like air-gapped or on-premise settings. You will act as the primary technical point of contact for your assigned accounts, with success measured by system functionality, client trust, and durable capability creation, rather than just ticket closure. This role offers significant autonomy and accountability for the system, client relationship, and overall outcomes.

Delhi onsite FullTime
DockerPythonRAG +1 more

Product Manager, API Infrastructure

5mo ago
OpenAI

OpenAI

We are seeking an experienced Product Manager to define and scale the construction of our data processing, data privacy, billing, and access controls products. You will set strategy and execute on projects like expanding our regional data processing footprint, enabling new inference caching controls in the API, or building APIs that make it easier for organizations to manage their spend limits. You will also define the strategy and ship foundational capabilities that ensure customers use OpenAI products securely, privately, and with enterprise-grade controls. This role partners deeply with engineering, security, legal, compliance, finance, and leadership to deliver high-trust, enterprise-grade systems.

San Francisco hybrid FullTime
OpenAIModel Serving

Principal ML Platform Engineer

5mo ago
s

synthesia.io

Synthesia is seeking a Principal Engineer to join their ML Platform team. This role focuses on building and operating the systems that enable researchers and product teams to train, serve, and deploy generative models efficiently and reliably. The team's work encompasses research infrastructure, production serving systems, internal tooling, and platform interfaces, with a growing emphasis on automation and agent-oriented workflows. This is a hands-on individual contributor position with significant ownership, where you will influence the evolution of the ML platform as it scales.

Europe remote FullTime
KubernetesPythonTerraform +1 more

Member of Technical Staff - Mid-Training Infra

5mo ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible. We build open models that empower individuals to control their intelligence and shape the future of AI. As a Member of Technical Staff focused on Mid-Training Infrastructure, you will be instrumental in designing, building, and operating large-scale GPU infrastructure crucial for high-throughput model inference and mid-training workloads. This role involves developing systems that support synthetic data generation and reinforcement learning pipelines at scale, as well as building high-performance inference platforms capable of serving and evaluating models across thousands of GPUs.

San Francisco, CA onsite FullTime
Reinforcement LearningModel Serving

Member of Technical Staff, Software Engineer

6mo ago
f

fireworks ai

Fireworks is seeking a Member of Technical Staff, Software Engineer to join their team. This role involves building the core backend systems that power Fireworks' platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. You will own major product surfaces from architecture to production, improving reliability, performance, and developer experience. This is platform engineering with product impact, where your systems will directly shape how customers build on top of AI. You will work closely with product, frontend, infra, and GTM to ship end-to-end features, and use AI tooling aggressively to automate tasks.

San Mateo hybrid FullTime
Fine-TuningPyTorchModel Serving +1 more

Audio Inference Engineer, Model Efficiency

10mo ago
Cohere

Cohere

Cohere is seeking an Audio Inference Engineer focused on Model Efficiency to join a fast-growing team of researchers and engineers. The mission of this team is to build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques. As an engineer on this team, you will advance core audio model serving metrics, including latency, throughput, and quality by diving deep into systems, identifying bottlenecks, and delivering creative solutions for audio processing and streaming workloads. You will collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real-time and streaming audio inference.

New York remote FullTime
CoherePythonC# +5 more

Member of Technical Staff, LLM Infrastructure

10mo ago
f

fireworks ai

As a Software Engineer on the AI Infrastructure team, you will help design the core systems that power Fireworks AI’s generative AI platform. You will build infrastructure and tools that ensure the reliability, performance, quality, and availability of our AI system. Your mission is to make Fireworks AI the most reliable and user-friendly generative AI platform in the world. You will partner closely with our cloud infrastructure, product, and performance teams to deliver infrastructure that bridges the gap between our customers and the ultra-performant proprietary Fireworks inference engine.

San Mateo hybrid FullTime
KubernetesPythonGo +5 more

Member of Technical Staff, Evals & Post-Training Product

10mo ago
f

fireworks ai

Fireworks is seeking a Member of Technical Staff, Evals & Post-Training Product to define how developers improve models on the Fireworks platform. This role combines scalable system design, deep data science, and model quality. You will build the infrastructure and workflows connecting evaluation and post-training, taking our evaluation setup to the next stage by improving programmatic access and scale. You will work across backend systems, sandbox infrastructure, and user-facing surfaces to simplify the authoring of evaluations, understanding of results, and rapid iteration.

San Mateo hybrid FullTime
Fine-TuningPyTorchModel Serving +1 more

Software Engineer - Model Products

11mo ago
B

Baseten

Baseten is seeking a Software Engineer to join their Model Performance team, focusing on the infrastructure that powers hosted API endpoints for cutting-edge open-source models. This role involves working on distributed systems, model serving, and developer experience to ensure models running on the Baseten platform are fast, reliable, and cost-efficient. You will contribute to defining how developers interact with AI models at scale, joining a high-impact team at the intersection of product, model performance, and infrastructure.

San Francisco hybrid FullTime
KubernetesSparkModel Serving

Machine Learning Infrastructure Engineer, Model Inference

1y ago
Abridge

Abridge

Abridge is seeking an ML Infrastructure Engineer, Model Inference to build and optimize the core inference infrastructure powering their machine learning models. This role is crucial for enhancing the scalability, efficiency, and performance of Abridge's AI-driven healthcare solutions. The engineer will collaborate with Infrastructure and Research teams to build, deploy, optimize, and orchestrate AI models, working on a platform that transforms patient-clinician conversations into structured clinical notes in real-time.

SF Office hybrid FullTime
KubernetesPyTorchTensorFlow +3 more

Applied Machine Learning Engineer

1y ago
f

fireworks ai

As an Applied Machine Learning Engineer, you will serve as a vital bridge between cutting-edge AI research and practical, real-world applications. Your work will focus on developing, fine-tuning, and operationalizing machine learning models that drive business value and enhance user experiences. This is a hands-on engineering role that combines deep technical expertise with a strong customer focus to deliver scalable AI solutions.

San Mateo hybrid FullTime
PythonFine-TuningPyTorch +4 more

Member of Technical Staff, Performance Optimization

1y ago
f

fireworks ai

Fireworks is seeking a Software Engineer focused on Performance Optimization to enhance the speed and efficiency of their AI infrastructure. This role involves optimizing performance across all levels of the technology stack, from low-level GPU kernels to large-scale distributed systems. The primary focus will be on maximizing the performance of demanding workloads such as large language models (LLMs), vision-language models (VLMs), and advanced video models. You will collaborate with research, infrastructure, and systems teams to identify and resolve performance bottlenecks, implement advanced optimizations, and scale AI systems for production use cases, directly influencing the speed, scalability, and cost-effectiveness of cutting-edge generative AI models.

San Mateo hybrid FullTime
KubernetesPyTorchModel Serving +1 more

Inference Engineer, Robotics

1y ago
OpenAI

OpenAI

We are seeking a GPU Inference Engineer to enhance model serving efficiency for our Robotics research. This high-impact role involves driving initiatives to optimize inference performance and scalability, as well as assisting researchers in developing inference-friendly models. This position is crucial for scaling the team's goals, enabling leadership to focus on higher-leverage initiatives by building a stronger technical foundation.

San Francisco hybrid FullTime
OpenAIModel Serving

Member of Technical Staff - Model Serving / API Backend Engineer

2y ago
B

Black Forest Labs

Black Forest Labs is at the forefront of generative AI, known for foundational technologies like Latent Diffusion and Stable Diffusion. We are building the next generation of creative tools used by millions worldwide. This role is crucial for bridging the gap between cutting-edge research and production-ready systems, ensuring that our advanced models can be efficiently deployed and experienced by users. You will be instrumental in accelerating the pace at which research breakthroughs become usable APIs and demos, directly impacting inference speed, API performance under load, and the overall user experience of our models.

$180k - $300k

San Francisco (United States) onsite FullTime
AWSAzureDocker +5 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.