Kubernetes Jobs

410 open roles mentioning Kubernetes

Technical Program Manager, Infrastructure

24d ago
Anthropic

Anthropic

Anthropic's Infrastructure organization is the engine that powers our mission, building and operating massive clusters for training frontier models, production infrastructure serving millions of users reliably, and developer platforms. As a Technical Program Manager for Infrastructure, you will coordinate complex programs with broad organizational impact, solving novel scaling challenges while maintaining security and reliability. This role is ideal for someone who thrives in ambiguity, makes others more effective, and partners closely with engineering leadership to drive strategic initiatives and ensure seamless coordination between research, engineering, and product teams.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicAWSAzure +3 more

Security Engineer, Detection & Response

24d ago
Anthropic

Anthropic

Anthropic is seeking an exceptional Detection and Response engineer to build solutions for monitoring threats, investigating incidents, and coordinating response efforts. This role offers the opportunity to shape security capabilities from the ground up alongside world-class research and security teams. The ideal candidate will be at the forefront of safeguarding advanced AI systems.

San Francisco, CA | New York City, NY | Seattle, WA; Washington, DC onsite
AnthropicKubernetesPython +1 more

Research Engineer, Discovery

24d ago
Anthropic

Anthropic

As a Research Engineer on our team, you will work end-to-end across the entire model stack, identifying and addressing key infrastructure blockers on the path to scientific AGI. You should have familiarity with elements of language model training, evaluation, and inference, and be eager to quickly dive into and get up to speed in areas where you are not yet an expert. This may include performance optimization, distributed systems, VM/sandboxing/container deployment, and large-scale data pipelines. Join us in our mission to develop advanced AI systems that push the frontiers of science and benefit humanity.

San Francisco, CA onsite
AnthropicAWSDocker +5 more

Senior Software Security Engineer

24d ago
Anthropic

Anthropic

Anthropic is building reliable, interpretable, and steerable AI systems to ensure AI is safe and beneficial for users and society. The Security Engineering team is crucial in protecting these AI systems and maintaining user trust. This role involves defining authentication architecture, designing cryptographic foundations for model weights and training data, and leading the developer security program. The team collaborates across identity and secrets management, developer security and supply chain, infrastructure security, and secure frameworks, with opportunities to contribute across these areas based on strengths and priorities.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicKubernetesPython +2 more

Research Engineer/Research Scientist, Pre-training

24d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer to join its Pre-training team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, contributing to the creation of safe, steerable, and trustworthy AI systems. The team is dedicated to ensuring that transformative AI systems are aligned with human interests and societal benefit.

Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY remote
AnthropicKubernetesPython +4 more

Research Engineer / Scientist, Alignment

24d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer/Scientist for its Alignment Science team. This role involves designing and executing machine learning experiments to understand and steer the behavior of advanced AI systems, with a focus on AI safety and potential risks from future human-level AI. You will collaborate with other teams on exploratory research, contributing to Anthropic's mission of creating reliable, interpretable, and steerable AI systems that are helpful, honest, and harmless.

San Francisco, CA onsite
AnthropicKubernetesPython +4 more

Staff Research Engineer, Discovery Team

24d ago
Anthropic

Anthropic

Anthropic is dedicated to building reliable, interpretable, and steerable AI systems that are safe and beneficial for society. As a Research Engineer on the Discovery Team, you will work end-to-end to identify and address key blockers on the path to scientific Artificial General Intelligence (AGI). This role involves improving models' abilities to use computers, acting as a laboratory for long-horizon tasks and a crucial component for scientific workflows. You will collaborate with a team of researchers and engineers focused on pushing the scientific frontier.

San Francisco, CA onsite
AnthropicDockerKubernetes +2 more

Technical Success Engineer

25d ago
L

Lambda

Lambda is seeking a Technical Success Engineer to join their Superintelligence business unit. This role is responsible for managing the deployment of AI cloud infrastructure for large, strategic customers, taking projects from contract signing to a fully operational production environment. You will act as the primary technical contact, ensuring customer requirements are met by validating configurations, identifying gaps, and collaborating with engineering and infrastructure teams to resolve issues. The goal is to ensure a seamless transition to production and customer self-sufficiency.

Remote, USA remote FullTime
KubernetesGoRAG

AI Software Engineer

25d ago
L

Legora

Legora is seeking software engineers to join its fast-moving team, working on technically difficult projects to revolutionize the legal industry with cutting-edge AI and LLM technology. The role involves architecting and implementing AI and LLM features end-to-end for their web platform, building systems for lawyers to work with complex documents using grounded, reliable AI, and designing/evaluating agentic workflows, retrieval strategies, and LLM-powered product experiences. This is a Stockholm-based, 5-day in-office role.

Stockholm HQ onsite FullTime
AWSAzureKubernetes +2 more

Member of Technical Staff, Infrastructure

25d ago
l

llamainndex ai

The Infra team at LlamaIndex is responsible for the foundational systems that power their AI product and the tools that enable engineers to develop, ship, and observe their code. They are seeking team members to design, build, and scale core infrastructure for a high-volume data platform for AI applications. The ideal candidate will have experience managing cloud infrastructure, navigating stages of scale, and empowering the engineering team. LlamaIndex values customer-obsession, collaboration, hard work, optimism, and ownership.

San Francisco hybrid FullTime
LlamaIndexKubernetesPython +3 more

Forward Deployed Engineer, Infrastructure Specialist (France)

26d ago
Cohere

Cohere

Cohere is seeking a Forward Deployed Engineer, Infrastructure Specialist to join our team in France. This role is crucial for deploying and managing our cutting-edge AI workspace platform, North, within enterprise environments. You will act as a key liaison between Cohere's core product and our clients' engineering teams, focusing on secure AI integration in critical sectors like finance and healthcare. The position offers a unique opportunity to solve complex problems at the forefront of Agentic AI and shape how enterprises leverage AI in real-world applications.

$20k - $40k

France hybrid FullTime
CohereAWSAzure +3 more

Forward Deployed Engineer (Training)

26d ago
B

Baseten

Baseten is seeking a Forward Deployed Engineer (Training) to work directly with leading AI companies, taking ownership of their technical outcomes on the Baseten platform. This role involves tackling complex challenges in serving and improving AI models at scale, spanning the entire model lifecycle from inference to post-training and the systems that connect them. You will act as a technical advisor, guiding customers from initial problem framing through to production deployment, and ensuring the quality and performance of their AI workloads through rigorous evaluation and optimization.

San Francisco hybrid FullTime
KubernetesPyTorchDeep Learning +2 more

Solutions Engineer (Chicago)

26d ago
L

Langchain

LangChain is seeking a Solutions Engineer to join their Deployed Engineering team. This role is a hands-on, highly technical position focused on partnering with account executives to help companies evaluate and implement LangChain's AI agent platform. You will own the technical win by scoping evaluations, designing proof-of-concepts, answering deep architecture questions, and ensuring successful production rollouts. This role sits at the intersection of engineering, product, and sales, with a fast feedback loop and measurable impact on closed deals and live deployments. You will work on challenging applied AI problems, directly shaping how AI agents are built and deployed in the real world.

$200k - $250k

Chicago, IL remote FullTime
LangGraphLangChainAWS +5 more

Frontend Engineer, Vision

27d ago
s

sarvam

Sarvam is building India's sovereign AI platform, focusing on research, models, infrastructure, and applications. This role is on the team responsible for building the serving harness for in-house vision-language models, turning them into production-ready document intelligence platforms. The goal is to match or exceed the performance of frontier hosted models at a fraction of the cost, handling the complexities of Indian documents at scale. You will be the frontend owner, working closely with backend and applied AI engineers to build the user interfaces that enable human trust and interaction with AI models.

Bengaluru onsite FullTime
KubernetesPythonTypeScript +4 more

Senior Backend Engineer, Vision

27d ago
s

sarvam

Sarvam is building India's full-stack sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. This role is for a Senior Backend Engineer who will own the architecture of the serving harness for Sarvam's vision models. The system needs to deliver frontier-grade extraction quality from in-house models at a national scale, while managing cost and latency budgets. This involves making critical trade-offs between accuracy, latency, and cost, and setting high standards for reliability, durability, and idempotency in document processing pipelines.

Bengaluru onsite FullTime
AzureKubernetesPython +4 more

Deployed Engineer (Early Career-NYC)

27d ago
L

Langchain

LangChain is seeking a Deployed Engineer to join their early-career team in NYC. This role focuses on working directly with customers to build and deploy production AI agents, transforming prototypes into reliable, scalable systems. You will partner with customer engineering teams across the full lifecycle, from pre-sales evaluations to post-deployment advisory work, ensuring customers successfully adopt innovative agentic engineering practices and ship reliable agents quickly. This is a hands-on, technical position at the intersection of engineering, product, and go-to-market, where your contributions will directly shape the real-world application of AI agents and influence the future direction of LangChain's platform.

$160k - $175k

New York, NY onsite FullTime
LangGraphLangChainAWS +5 more

Deployed Engineer (Early Career- SF)

27d ago
L

Langchain

This role is for an early-career engineer focused on helping customers adopt and deploy AI agents in production. You will work directly with customer engineering teams to co-architect and build AI agent systems, ensuring they are reliable and scalable. The position involves a blend of pre-sales technical wins, post-deployment advisory, and providing feedback to the product team. You will be instrumental in shaping how AI agents are built and used in the real world, working on challenging problems in applied AI with a fast feedback loop and visible impact.

$160k - $175k

San Francisco, CA onsite FullTime
LangGraphLangChainAWS +5 more

Detection and Response Engineer

28d ago
M

Modal

AI needs a new infrastructure layer, and Modal is building it. We are seeking a Detection & Response Engineer to develop systems for identifying, investigating, and responding to threats across our platform. This is an engineering role focused on automation, where you will build scalable detections, investigation tooling, and response capabilities, leveraging AI to enhance signal, investigation speed, and operational effectiveness. You will collaborate closely with infrastructure, platform, and security engineers to ensure every incident contributes to platform resilience.

New York onsite FullTime
KubernetesSQL

Security Engineer - Detection & Response

28d ago
L

Langchain

LangChain is seeking a hands-on Detection & Response Engineer to join their Security team. This role is crucial for monitoring, preventing, and learning from threats across LangChain's production platform, cloud infrastructure, and agentic workloads. You will collaborate with Product Security to translate threat models and incidents into effective telemetry and detections. The primary focus is on engineering scalable detection, investigation, and response systems that enhance the Security team's impact beyond manual operations, requiring strong software engineering skills and a builder's mindset.

$180k - $240k

San Francisco, CA onsite FullTime
LangGraphLangChainAWS +5 more

Member of Technical Staff - Reliability Engineering

28d ago
f

fireworks ai

Fireworks AI is seeking a Member of Technical Staff focused on Reliability Engineering to ensure the dependable operation of their AI platform. This role involves working across cloud infrastructure, AI systems, and product teams to guarantee seamless integration, graceful failure handling, and robust performance under load. You will be instrumental in defining reliability standards, owning the reliability toolchain, and ensuring a positive customer experience by addressing failures that span across systems. The position requires a proactive approach to incident management, a commitment to reducing operational toil through automation, and strong collaboration with various engineering teams.

San Mateo hybrid FullTime
DockerKubernetesPython +5 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.