Model Serving Jobs
74 open roles mentioning Model Serving
Strategic Deployment Engineer, Chanakya
sarvam
Sarvam is building India's full-stack sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. As a Strategic Deployment Engineer, you will be embedded with clients, owning the entire lifecycle of AI system deployments in complex and often constrained environments like air-gapped or on-premise settings. You will act as the primary technical point of contact for your assigned accounts, with success measured by system functionality, client trust, and durable capability creation, rather than just ticket closure. This role offers significant autonomy and accountability for the system, client relationship, and overall outcomes.
Product Manager, API Infrastructure
OpenAI
We are seeking an experienced Product Manager to define and scale the construction of our data processing, data privacy, billing, and access controls products. You will set strategy and execute on projects like expanding our regional data processing footprint, enabling new inference caching controls in the API, or building APIs that make it easier for organizations to manage their spend limits. You will also define the strategy and ship foundational capabilities that ensure customers use OpenAI products securely, privately, and with enterprise-grade controls. This role partners deeply with engineering, security, legal, compliance, finance, and leadership to deliver high-trust, enterprise-grade systems.
Principal ML Platform Engineer
synthesia.io
Synthesia is seeking a Principal Engineer to join their ML Platform team. This role focuses on building and operating the systems that enable researchers and product teams to train, serve, and deploy generative models efficiently and reliably. The team's work encompasses research infrastructure, production serving systems, internal tooling, and platform interfaces, with a growing emphasis on automation and agent-oriented workflows. This is a hands-on individual contributor position with significant ownership, where you will influence the evolution of the ML platform as it scales.
Member of Technical Staff - Mid-Training Infra
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible. We build open models that empower individuals to control their intelligence and shape the future of AI. As a Member of Technical Staff focused on Mid-Training Infrastructure, you will be instrumental in designing, building, and operating large-scale GPU infrastructure crucial for high-throughput model inference and mid-training workloads. This role involves developing systems that support synthetic data generation and reinforcement learning pipelines at scale, as well as building high-performance inference platforms capable of serving and evaluating models across thousands of GPUs.
Member of Technical Staff, Software Engineer
fireworks ai
Fireworks is seeking a Member of Technical Staff, Software Engineer to join their team. This role involves building the core backend systems that power Fireworks' platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. You will own major product surfaces from architecture to production, improving reliability, performance, and developer experience. This is platform engineering with product impact, where your systems will directly shape how customers build on top of AI. You will work closely with product, frontend, infra, and GTM to ship end-to-end features, and use AI tooling aggressively to automate tasks.
Audio Inference Engineer, Model Efficiency
Cohere
Cohere is seeking an Audio Inference Engineer focused on Model Efficiency to join a fast-growing team of researchers and engineers. The mission of this team is to build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques. As an engineer on this team, you will advance core audio model serving metrics, including latency, throughput, and quality by diving deep into systems, identifying bottlenecks, and delivering creative solutions for audio processing and streaming workloads. You will collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real-time and streaming audio inference.
Member of Technical Staff, LLM Infrastructure
fireworks ai
As a Software Engineer on the AI Infrastructure team, you will help design the core systems that power Fireworks AI’s generative AI platform. You will build infrastructure and tools that ensure the reliability, performance, quality, and availability of our AI system. Your mission is to make Fireworks AI the most reliable and user-friendly generative AI platform in the world. You will partner closely with our cloud infrastructure, product, and performance teams to deliver infrastructure that bridges the gap between our customers and the ultra-performant proprietary Fireworks inference engine.
Member of Technical Staff, Evals & Post-Training Product
fireworks ai
Fireworks is seeking a Member of Technical Staff, Evals & Post-Training Product to define how developers improve models on the Fireworks platform. This role combines scalable system design, deep data science, and model quality. You will build the infrastructure and workflows connecting evaluation and post-training, taking our evaluation setup to the next stage by improving programmatic access and scale. You will work across backend systems, sandbox infrastructure, and user-facing surfaces to simplify the authoring of evaluations, understanding of results, and rapid iteration.
Software Engineer - Model Products
Baseten
Baseten is seeking a Software Engineer to join their Model Performance team, focusing on the infrastructure that powers hosted API endpoints for cutting-edge open-source models. This role involves working on distributed systems, model serving, and developer experience to ensure models running on the Baseten platform are fast, reliable, and cost-efficient. You will contribute to defining how developers interact with AI models at scale, joining a high-impact team at the intersection of product, model performance, and infrastructure.
Machine Learning Infrastructure Engineer, Model Inference
Abridge
Abridge is seeking an ML Infrastructure Engineer, Model Inference to build and optimize the core inference infrastructure powering their machine learning models. This role is crucial for enhancing the scalability, efficiency, and performance of Abridge's AI-driven healthcare solutions. The engineer will collaborate with Infrastructure and Research teams to build, deploy, optimize, and orchestrate AI models, working on a platform that transforms patient-clinician conversations into structured clinical notes in real-time.
Applied Machine Learning Engineer
fireworks ai
As an Applied Machine Learning Engineer, you will serve as a vital bridge between cutting-edge AI research and practical, real-world applications. Your work will focus on developing, fine-tuning, and operationalizing machine learning models that drive business value and enhance user experiences. This is a hands-on engineering role that combines deep technical expertise with a strong customer focus to deliver scalable AI solutions.
Member of Technical Staff, Performance Optimization
fireworks ai
Fireworks is seeking a Software Engineer focused on Performance Optimization to enhance the speed and efficiency of their AI infrastructure. This role involves optimizing performance across all levels of the technology stack, from low-level GPU kernels to large-scale distributed systems. The primary focus will be on maximizing the performance of demanding workloads such as large language models (LLMs), vision-language models (VLMs), and advanced video models. You will collaborate with research, infrastructure, and systems teams to identify and resolve performance bottlenecks, implement advanced optimizations, and scale AI systems for production use cases, directly influencing the speed, scalability, and cost-effectiveness of cutting-edge generative AI models.
Inference Engineer, Robotics
OpenAI
We are seeking a GPU Inference Engineer to enhance model serving efficiency for our Robotics research. This high-impact role involves driving initiatives to optimize inference performance and scalability, as well as assisting researchers in developing inference-friendly models. This position is crucial for scaling the team's goals, enabling leadership to focus on higher-leverage initiatives by building a stronger technical foundation.
Member of Technical Staff - Model Serving / API Backend Engineer
Black Forest Labs
Black Forest Labs is at the forefront of generative AI, known for foundational technologies like Latent Diffusion and Stable Diffusion. We are building the next generation of creative tools used by millions worldwide. This role is crucial for bridging the gap between cutting-edge research and production-ready systems, ensuring that our advanced models can be efficiently deployed and experienced by users. You will be instrumental in accelerating the pace at which research breakthroughs become usable APIs and demos, directly impacting inference speed, API performance under load, and the overall user experience of our models.
$180k - $300k