Kubernetes Jobs

204 open roles mentioning Kubernetes

Senior Software Engineer, Production Engineering

11h ago
Harvey

Harvey

Harvey is seeking a Production Engineer to help build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth. The ideal candidate will have a systems-thinking mindset and a passion for building simple, reliable, and scalable systems.

$161k - $242k

New York remote FullTime
AWSAzureKubernetes +7 more

Staff Software Engineer, Production Engineering

11h ago
Harvey

Harvey

Harvey is seeking a Production Engineer to build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth.

$231k - $340k

New York remote FullTime
AWSAzureKubernetes +7 more

Staff Software Engineer, Production Engineering

11h ago
Harvey

Harvey

Harvey is seeking a Production Engineer to build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth.

$231k - $340k

San Francisco remote FullTime
AWSAzureKubernetes +7 more

Senior Software Engineer, Production Engineering

11h ago
Harvey

Harvey

Harvey is seeking a Production Engineer to help build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth. The ideal candidate will have a systems-thinking mindset and a passion for building simple, reliable, and scalable systems.

$161k - $242k

San Francisco remote FullTime
AWSAzureKubernetes +7 more

Infrastructure engineer (UK)

14h ago
W

Writer

At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need.

London, UK hybrid FullTime
AWSAzureKubernetes +9 more

Infrastructure engineer

14h ago
W

Writer

At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need.

New York City, NY hybrid FullTime
AWSAzureKubernetes +10 more

Engineering Manager, Model Infrastructure

9d ago
Harvey

Harvey

Harvey is transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. This role offers a rare chance to help build a generational company at a true inflection point, with strong product-market fit and world-class investor support. The team moves fast, takes ownership, and is deeply committed to the mission, operating with intensity and pushing for excellence. As the Engineering Manager for Model Infrastructure, you will lead the team responsible for the platform powering every model request across Harvey, partnering closely with AI Research, Product Engineering, and Infrastructure to ensure reliability, scalability, and cost-efficiency. This is a strategic engineering organization, critical for every product capability, and will evolve to build the infrastructure for Harvey to train, evaluate, deploy, and operate its own frontier AI models.

$260k - $340k

San Francisco hybrid FullTime
OpenAIAnthropicAzure +5 more

Engineering Manager, GPU Infrastructure

9d ago
Cohere

Cohere

Cohere is a leading enterprise AI company building cutting-edge foundation AI models and end-to-end products. The GPU Clusters team is central to Cohere's infrastructure, responsible for building and operating the superclusters that power our frontier AI models. This role involves enabling research and development at the intersection of cutting-edge hardware, distributed systems, and AI research. As an Engineering Manager, you will lead a team of highly motivated engineers passionate about GPU infrastructure and AI, fostering a culture of technical excellence and innovation in a remote-first environment. This is a unique opportunity to shape the infrastructure powering the next generation of AI.

United States onsite FullTime
CohereKubernetesPyTorch +4 more

Staff+ Software Engineer, Platform

13d ago
Anthropic

Anthropic

Anthropic is seeking experienced software engineers to join its Platform organization. This team builds the foundational primitives that accelerate product development and owns the infrastructure and systems that enable reliable, scaled product shipping for internal and external users. You will independently scope complex, multi-month projects, drive cross-organizational alignment, and make architectural decisions that shape how Anthropic builds and scales its products. You will partner directly with research to productize cutting-edge capabilities, impacting the platform used by hundreds of thousands of companies and engineers daily.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicAWSKubernetes +5 more

Software Engineer, Developer Productivity, AI Tools

13d ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking a developer productivity engineer to enhance internal software development processes, focusing on safety, speed, and user experience. This role will concentrate on AI tools and coding agents, collaborating with platform, security, and product engineers to build cutting-edge tooling for AI-assisted software development and significantly accelerate the inner development loop. The position involves both establishing company-wide platforms and assisting individual developers in optimizing their workflows.

$350k - $475k

San Francisco onsite
OpenAIMistralDocker +5 more

Site Reliability Engineer (SRE)

13d ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking a Site Reliability Engineer (SRE) to ensure the end-to-end reliability of their Tinker platform. This role involves working closely with engineers and research teams to enhance the robustness and resilience of every system layer. The SRE will be instrumental in maintaining and improving the infrastructure that supports custom AI model fine-tuning, ensuring a seamless experience for researchers and developers.

$350k - $475k

San Francisco onsite
OpenAIMistralKubernetes +3 more

Staff+ Software Engineer, Capacity Engineering

13d ago
Anthropic

Anthropic

Anthropic is building reliable, interpretable, and steerable AI systems to ensure AI is safe and beneficial. The Capacity Engineering team plays a crucial role in managing one of the industry's largest and fastest-growing infrastructure fleets. This team is responsible for ensuring all infrastructure resources are accounted for, well-utilized, and efficiently allocated. As an engineer on this team, you will develop the production systems that power these efforts, including data pipelines for telemetry ingestion, observability tools for real-time fleet health visibility, and performance instrumentation to measure hardware utilization across various workloads. You will write production-quality code, operate at scale with Kubernetes-native infrastructure, and significantly influence decisions regarding Anthropic's substantial infrastructure spend. This role involves close collaboration with research engineering, infrastructure, inference, and finance teams, requiring comfort across data engineering, systems engineering, and observability in a high-autonomy environment.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicAWSAzure +4 more

Research Engineer, Infrastructure, Inference

14d ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking an infrastructure research engineer to design, optimize, and scale the systems that power large AI models. The goal is to make inference faster, more cost-effective, more reliable, and more reproducible, enabling research teams to focus on advancing model capabilities. This role is crucial for ensuring that every experiment, evaluation, and deployment runs smoothly at scale, with a focus on performant and efficient model inference for both real-world applications and research acceleration.

$350k - $475k

San Francisco onsite
OpenAIMistralKubernetes +3 more

Senior Software Engineer, Inference

14d ago
Anthropic

Anthropic

Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry's largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team has a dual mandate: maximizing compute efficiency to serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms.

London, UK onsite
AnthropicAWSAzure +5 more

Staff Software Engineer, Inference

14d ago
Anthropic

Anthropic

Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry's largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team has a dual mandate: maximizing compute efficiency to serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms. As a Staff Software Engineer on our Inference team, you will work end to end, identifying and addressing key infrastructure blockers to serve Claude to millions of users while enabling breakthrough AI research.

London, UK onsite
AnthropicAWSAzure +5 more

Member of Technical Staff - Engineering Lead, Platform Foundations

16d ago
Reflection ai

Reflection ai

Reflection is seeking a Platform Foundations Lead to build and operate the foundational layer that all engineering teams depend on. This includes cloud and multi-cloud infrastructure, networking, security, and developer tooling. The role involves leading a team of infrastructure engineers, guiding technical and architectural decisions, and working closely with other engineering teams. You will also manage cloud and infrastructure vendors and contribute as an individual contributor to stay close to the technical stack.

San Francisco, CA onsite FullTime
KubernetesTerraform

Manager of Technology & Security Engineering

16d ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible. We build open models that empower individuals to control their intelligence and shape the future of AI. The Head of Technology & Security Engineering will architect and operate the security engineering foundation protecting our corporate environment, multi-cloud research infrastructure, and significant GPU training capacity. This leadership role encompasses the full technical security stack, from end-user compute and zero-trust access to cloud security, detection engineering, and security-as-code, enabling rapid research progress without compromising rigor. Positioned at the intersection of security engineering, infrastructure, and research operations, this role is crucial for safeguarding sensitive assets like model weights, training data, and GPU capacity, while maintaining a low-friction environment for researchers and engineers. The ideal candidate possesses deep technical expertise and executive presence to lead the security engineering function, represent the company's security posture, and manage vendor and contractual risks.

New York, NY onsite FullTime
KubernetesPythonTerraform

Staff + Sr. Software Engineer, Cloud Inference Launch Engineering

16d ago
Anthropic

Anthropic

Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. The Cloud Inference team is responsible for scaling and optimizing Claude to serve massive audiences across AWS, GCP, Azure, and other cloud providers. This involves end-to-end ownership of Claude on each platform, from API integration and request routing to inference execution and capacity management. The model & inference launch team specifically focuses on the validation pipeline for the inference server and load balancer, ensuring every inference change, including model launches and performance improvements, is deployed with correctness, performance, and reliability. This high-leverage infrastructure work is critical for fast and cost-effective deployment of frontier models and features, directly impacting compute capacity and the speed at which innovations reach production.

San Francisco, CA onsite
AnthropicAWSAzure +5 more

Staff Software Engineer, Inference

16d ago
Anthropic

Anthropic

Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry's largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team has a dual mandate: maximizing compute efficiency to serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms.

Dublin, IE onsite
AnthropicAWSKubernetes +4 more

Technical Program Manager, Infrastructure

16d ago
Anthropic

Anthropic

Anthropic's Infrastructure organization is the engine that powers our mission, building and operating massive clusters for training frontier models, production infrastructure serving millions of users reliably, and developer platforms. As a Technical Program Manager for Infrastructure, you will coordinate complex programs with broad organizational impact, solving novel scaling challenges while maintaining security and reliability. This role is ideal for someone who thrives in ambiguity, makes others more effective, and partners closely with engineering leadership to drive strategic initiatives and ensure seamless coordination between research, engineering, and product teams.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicAWSAzure +3 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.