Kubernetes Jobs

410 open roles mentioning Kubernetes

Lead Member of Technical Staff, Inference Infrastructure

4mo ago
Cohere

Cohere

Cohere is seeking a Lead Member of Technical Staff to join the Model Serving team. This role will provide technical leadership across multiple teams, driving the architecture and strategy for deploying optimized NLP models to production in low latency, high throughput, and high availability environments. You will be responsible for developing, deploying, and operating the AI platform that delivers Cohere's large language models through easy-to-use API endpoints. The ideal candidate will have a strong background in leading the design of high-performance, scalable, and reliable machine learning systems and will serve as a key point of contact for customers, leading the design of customized deployments and mentoring engineers.

San Francisco remote FullTime
CohereAWSAzure +5 more

Staff Software Engineer, Core Infrastructure

4mo ago
Harvey

Harvey

Harvey is transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. This is a rare chance to help build a generational company at a true inflection point, scaling fast and defining a new category. The work is ambitious, the bar is high, and the opportunity for growth is unmatched. As a Staff Software Engineer on the Core Infrastructure team, you will play a critical role in designing and building new infrastructure systems while scaling and strengthening existing ones. Our infrastructure powers every user interaction with Harvey, processing billions of prompt tokens and millions of daily requests across our global legal AI platform. You'll work in an environment balanced between innovation and operational excellence, ensuring Harvey remains resilient and efficient as it scales products, regions, customers, and usage. Your contributions will directly impact the reliability, scalability, and security of our platform.

$201k - $264k

New York onsite FullTime
AWSAzureKubernetes +5 more

Staff Software Engineer, Core Infrastructure

4mo ago
Harvey

Harvey

Harvey is transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. This is a rare chance to help build a generational company at a true inflection point, scaling fast and defining a new category. The work is ambitious, the bar is high, and the opportunity for growth is unmatched. The team moves fast, takes ownership, and is deeply committed to the mission, operating with intensity and pushing for excellence. As a Staff Software Engineer on the Core Infrastructure team, you will play a critical role in designing and building new infrastructure systems while scaling and strengthening existing ones. Your contributions will directly impact the reliability, scalability, and security of our platform as we serve the world's leading law firms and professional service providers.

$236k - $290k

San Francisco onsite FullTime
AWSAzureKubernetes +5 more

Software Engineer - Voice AI (Inference Runtime)

4mo ago
B

Baseten

Baseten is seeking a highly impactful individual to lead the Voice AI product area, owning the end-to-end development and implementation of their in-house inference stack for Voice AI models. This role involves partnering closely with various engineering teams to push the boundaries of Voice AI, making a significant impact on industries like productivity, customer service, and education. You will be responsible for bringing state-of-the-art open-source models into production, focusing on optimizing model serving for latency, throughput, and GPU efficiency, and building large-scale, real-time infrastructure for multi-model voice agents.

San Francisco hybrid FullTime
DockerKubernetesPython +5 more

Research Engineer, Data Infrastructure

4mo ago
Mistral AI

Mistral AI

Mistral AI is seeking a Research Engineer focused on Data Infrastructure to build and operate the next generation of our data systems. This role involves designing and scaling massive compute fleets and storage systems for high performance and scalability. You will contribute to a future of decoupled control and data planes, scaling big data compute and storage platforms while ensuring secure and governed data access for MLOps and research. The position requires full lifecycle ownership, from architecting migrations away from legacy orchestrators to implementing production-grade pipelines and participating in on-call rotations for critical training jobs.

Palo Alto remote Full-time
MistralKubernetesPython +5 more

Engineering Manager, FDE Infrastructure (UK)

4mo ago
Cohere

Cohere

Cohere is seeking an Engineering Manager to lead their Deployment Engineering team in EMEA. This is a hands-on leadership role where you will manage a team of Forward Deployed Engineers responsible for deploying Cohere's North platform into customer environments. You will drive end-to-end deployments, take ownership of customer technical implementation, and collaborate with Product, Engineering, and Sales teams. The role requires mentoring the team on cloud infrastructure, Kubernetes, and enterprise-grade deployments, as well as optimizing performance and defining scaling guidelines for compute resources.

United Kingdom remote FullTime
CohereAWSAzure +2 more

Infrastructure Security Engineer

4mo ago
M

Modal

AI needs a new infrastructure layer, and this role is central to building it at Modal. We are seeking an Infrastructure Security Engineer to design and secure the core systems powering our platform. This position emphasizes building security directly into our infrastructure, covering aspects like container isolation, orchestration, identity, and secrets management within a multi-tenant, cloud-native environment. You will collaborate closely with engineering teams to establish secure primitives and ensure the platform's resilience, scalability, and trustworthiness by design. This is a hands-on, deeply technical role focused on practical implementation rather than compliance or policy.

New York onsite FullTime
AWSKubernetesGCP

ML Ops Engineer, Chanakya

4mo ago
s

sarvam

Sarvam is building India's full-stack sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. The MLOps Engineer will own the model lifecycle across all deployments, ensuring systems are always operational, accurate, and auditable. This role involves supporting field engineers and managing deployment infrastructure for new products, with an uncompromising standard for reliability, as model failures are considered operational risks.

Delhi onsite FullTime
DockerKubernetesPython +3 more

Embedded Infrastructure Engineer, Chanakya

5mo ago
s

sarvam

Embedded Infrastructure Engineers at Sarvam design, build, and maintain the data infrastructure essential for deploying AI systems at client sites. You will collaborate with Embedded Data Scientists and Strategic Deployment Engineers to ensure the reliable and performant ingestion, storage, querying, and serving of terabyte-scale datasets to AI reasoning engines. This involves constructing and managing data platforms capable of handling large volumes of structured records, documents, imagery, audio, and geospatial data, including databases, object stores, ingestion pipelines, and processing layers. You will be responsible for making key decisions regarding storage architecture, indexing strategies, pipeline orchestration, and system performance, often in challenging environments like air-gapped or operationally sensitive settings that preclude the use of standard cloud services or enterprise tooling. Ultimately, you will own the reliability and performance of the infrastructure layer for your assigned accounts.

Delhi onsite FullTime
WeaviateQdrantMilvus +5 more

Senior Full-Stack Engineer

5mo ago
N

Nabla

Nabla is seeking a Senior Full-Stack Engineer to join a cross-functional squad and contribute end-to-end to a specific part of the product. This role involves taking a leading role in building and scaling Nabla’s AI-powered healthcare platform, working across the stack from back-end services to web, desktop, and mobile applications. You will collaborate closely with Product, Design, and Machine Learning teams to deliver high-impact features that directly serve healthcare professionals, with the goal of restoring the human connection at the heart of healthcare by streamlining clinical documentation.

Paris office onsite FullTime
KubernetesPythonTypeScript +5 more

Security Engineer, Cloud Infrastructure

5mo ago
M

Mercor

Mercor is seeking a Security Engineer to own cloud and infrastructure security at a company where tenant isolation is a critical enterprise requirement. You will architect multi-account AWS isolation, harden Kubernetes clusters, deploy cloud security posture management, and build the infrastructure that lets Mercor serve enterprise clients who demand the highest security bar. This role involves using AI heavily in security work, including building alongside AI code-gen tools, using LLMs to accelerate infrastructure review and policy authoring, and automating repetitive tasks. If you prefer writing Terraform modules over filling out spreadsheets, you will fit in here.

San Francisco or NYC onsite FullTime
AWSKubernetesTerraform

Deployed Engineer (Raleigh)

5mo ago
L

Langchain

LangChain is seeking a Deployed Engineer to join their team, focusing on making intelligent agents ubiquitous. This role involves working on challenging applied AI problems, building production systems that real teams depend on, and directly shaping how AI agents are built and used in the real world. The feedback loop is fast, the impact is visible, and the work contributes to the evolution of AI agent technology.

$150k - $250k

Raleigh, NC remote FullTime
LangGraphLangChainAWS +5 more

Deployed Engineer (Charlotte)

5mo ago
L

Langchain

LangChain is seeking a Deployed Engineer to join their team and work on challenging applied AI problems. This role focuses on building and deploying production AI agents that real teams depend on, offering a fast feedback loop and visible impact. You will be instrumental in shaping how AI agents are built and adopted in the real world, moving beyond demos and research to create reliable systems.

$150k - $250k

Charlotte, NC remote FullTime
LangGraphLangChainAWS +5 more

Member of Technical Staff (AI Infrastructure Engineer)

5mo ago
P

Perplexity AI

We are seeking an AI Infrastructure Engineer to join our expanding team. In this role, you will collaborate closely with our Inference and Research teams to construct, deploy, and enhance our extensive AI training and inference clusters. Your work will involve managing Kubernetes and Slurm environments, optimizing distributed training for large language models, and developing robust orchestration systems. This position offers the opportunity to significantly impact the performance and scalability of our AI infrastructure.

London hybrid FullTime
AWSKubernetesPython +5 more

Member of Technical Staff (AI Infrastructure Engineer)

5mo ago
P

Perplexity AI

We are seeking an AI Infrastructure Engineer to join our expanding team. In this role, you will collaborate closely with our Inference and Research teams to construct, deploy, and enhance our large-scale AI training and inference clusters. Our work involves Kubernetes, Slurm, Python, C++, PyTorch, and primarily operates on AWS.

San Francisco onsite FullTime
AWSKubernetesPython +5 more

Member of Technical Staff (AI Inference Engineer)

5mo ago
P

Perplexity AI

We are seeking an engineer to join our team responsible for building and running the inference engine behind Perplexity's queries. This role involves deploying dozens of model architectures at scale, managing tight latency and cost budgets, and working with a stack including Rust, Python, CUDA, and CuTe DSL. You will contribute to supporting new models, migrating GPU kernels, developing a Rust-native serving runtime, optimizing performance, and enhancing reliability and observability.

San Francisco onsite FullTime
KubernetesPythonRust +4 more

Member of Technical Staff (AI Inference Engineer)

5mo ago
P

Perplexity AI

We are seeking an AI Inference Engineer to join our dynamic team. This role is central to Perplexity's operations, as you will build and manage the inference engine that powers every query. You will deploy a variety of model architectures at scale, focusing on meeting stringent latency and cost requirements. Our technology stack includes Rust, Python, CUDA, and the CuTe DSL.

London onsite FullTime
KubernetesPythonRust +4 more

Mistral Cloud - Software Engineer, Backend (Golang)

5mo ago
Mistral AI

Mistral AI

Mistral AI is dedicated to simplifying tasks, saving time, and enhancing learning and creativity through AI. We democratize AI with high-performance, open-source models and a comprehensive AI platform for enterprise needs. We are seeking passionate and talented software engineers to join our team and play a key role in developing the core systems that drive our Cloud solutions. Your work will involve building robust, high-performance backend systems and APIs to support our rapidly expanding enterprise user base.

Paris remote Full-time
MistralAWSAzure +11 more

Research Engineer, Data Infrastructure

5mo ago
Mistral AI

Mistral AI

Mistral AI is seeking a Research Engineer focused on Data Infrastructure to architect and build the backbone of our frontier model training and fine-tuning ecosystem. This role involves designing and scaling massive compute fleets and storage systems for high performance and scalability, with a vision towards exabyte-scale architecture. You will contribute to a strategic transition from legacy scheduling to modern orchestration, implementing sophisticated multi-cluster orchestration and cloud-bursting capabilities. The position requires full lifecycle ownership, from architecting migrations away from legacy orchestrators to implementing production-grade pipelines and participating in on-call rotations for critical training jobs.

Paris hybrid Full-time
MistralKubernetesPython +5 more

Principal ML Platform Engineer

5mo ago
s

synthesia.io

Synthesia is seeking a Principal Engineer to join their ML Platform team. This role focuses on building and operating the systems that enable researchers and product teams to train, serve, and deploy generative models efficiently and reliably. The team's work encompasses research infrastructure, production serving systems, internal tooling, and platform interfaces, with a growing emphasis on automation and agent-oriented workflows. This is a hands-on individual contributor position with significant ownership, where you will influence the evolution of the ML platform as it scales.

Europe remote FullTime
KubernetesPythonTerraform +1 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.