Kubernetes Jobs

411 open roles mentioning Kubernetes

Senior Infrastructure Security Engineer

1mo ago
c

crusoe

Crusoe is seeking a Senior Infrastructure Security Engineer to establish a high-assurance security posture through deep visibility and granular access controls. This role involves leveraging advanced tooling like Wiz to maintain an uncompromising security posture, using infrastructure service building as the primary mechanism to enforce these standards across our global compute platform. You will be instrumental in building the future of AI infrastructure with an energy-first approach.

Dublin - IE onsite FullTime
AWSKubernetesPython +3 more

AI Field Engineer, Singapore

1mo ago
f

fireworks ai

Fireworks is seeking an AI Field Engineer to join their team in Singapore. This role is at the forefront of technical engagement, embedding with key customers and partners to transform complex AI challenges into production-ready systems. You will operate at the nexus of engineering, product development, and customer delivery, actively building proofs-of-concept (POCs), minimum viable products (MVPs), and production integrations. Simultaneously, you will engage in high-level discussions regarding architecture, strategy, and business impact. The position requires a blend of hands-on coding, performance benchmarking, debugging, and deployment architecture, alongside leading customer discovery, aligning stakeholders, and translating customer needs into product enhancements.

Singapore onsite FullTime
AWSAzureKubernetes +5 more

AI Field Engineer, EMEA

1mo ago
f

fireworks ai

Fireworks is seeking an AI Field Engineer to join our team in EMEA. This role is at the forefront of our technical engagement with ambitious customers and technology partners, focusing on transforming complex AI challenges into production-ready systems. You will operate at the intersection of engineering, product, and customer delivery, actively building Proofs of Concepts (POCs), Minimum Viable Products (MVPs), and production integrations. This position requires a blend of hands-on technical execution and the ability to engage in executive-level discussions regarding architecture, strategy, and business outcomes. The role emphasizes building strong relationships and trust through in-person interactions with clients, particularly within large organizations and digital-native companies adopting GenAI.

London onsite FullTime
AWSAzureKubernetes +5 more

Engineering Manager, Model Infrastructure

1mo ago
Harvey

Harvey

Harvey is transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. This role offers a rare chance to help build a generational company at a true inflection point, with strong product-market fit and world-class investor support. The team moves fast, takes ownership, and is deeply committed to the mission, operating with intensity and pushing for excellence. As the Engineering Manager for Model Infrastructure, you will lead the team responsible for the platform powering every model request across Harvey, partnering closely with AI Research, Product Engineering, and Infrastructure to ensure reliability, scalability, and cost-efficiency. This is a strategic engineering organization, critical for every product capability, and will evolve to build the infrastructure for Harvey to train, evaluate, deploy, and operate its own frontier AI models.

$260k - $340k

San Francisco hybrid FullTime
OpenAIAnthropicAzure +5 more

Senior Product Manager, Orchestration

1mo ago
c

crusoe

Crusoe is seeking a Senior Product Manager for Orchestration to own major product areas within their Managed Orchestration portfolio. This role involves end-to-end ownership of Kubernetes cluster lifecycle, control-plane capabilities, node health, managed add-ons, workload scheduling, Slurm on Kubernetes, and customer-facing APIs for large-scale AI infrastructure. It's a deeply technical and hands-on position requiring close collaboration with engineering to translate complex infrastructure challenges into clear, durable product designs, focusing on detailed behavior, edge cases, failure modes, and thoughtful trade-offs to ensure product quality and customer satisfaction.

$170k - $205k

San Francisco, CA - US onsite FullTime
KubernetesFine-Tuning

Staff Applied AI Inference Engineer

1mo ago
c

crusoe

Crusoe is seeking a Staff Applied AI Inference Engineer to accelerate the abundance of energy and intelligence by optimizing large language models for production environments. This role involves owning the inference stack end-to-end, from profiling costs and implementing modern optimization techniques to deep dives into serving code when defaults are insufficient. The work is applied, focusing on real customer deployments with varying models, traffic, latency targets, and cost constraints. You will collaborate with customer engineering teams to tailor deployments, transition workloads from proof-of-concept to production, and ensure engineered gains are realized by users. This is a hands-on engineering position requiring coding, profiling, and low-level optimization, with a customer-facing component involving product and technical solutions work.

$215k - $260k

San Francisco, CA - US onsite FullTime
DockerKubernetesPython +2 more

Software Engineer, Developer Productivity, AI Tools

1mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking a developer productivity engineer to enhance internal software development processes, focusing on safety, speed, and user experience. This role will concentrate on AI tools and coding agents, collaborating with platform, security, and product engineers to build cutting-edge tooling for AI-assisted software development and significantly accelerate the inner development loop. The position involves both establishing company-wide platforms and assisting individual developers in optimizing their workflows.

$350k - $475k

San Francisco onsite
OpenAIMistralDocker +5 more

Site Reliability Engineer (SRE)

1mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking a Site Reliability Engineer (SRE) to ensure the end-to-end reliability of their Tinker platform. This role involves working closely with engineers and research teams to enhance the robustness and resilience of every system layer. The SRE will be instrumental in maintaining and improving the infrastructure that supports custom AI model fine-tuning, ensuring a seamless experience for researchers and developers.

$350k - $475k

San Francisco onsite
OpenAIMistralKubernetes +3 more

Sr. Cloud Support Engineer - Weekend Shift

1mo ago
c

crusoe

Crusoe Cloud is revolutionizing high-performance computing by offering sustainable, low-cost GPU compute power. As a Cloud Support Engineer, you will be the primary point of contact for technical support, ensuring our customers can seamlessly utilize Crusoe Cloud for groundbreaking advancements in fields like AI/ML, physics simulations, and computational biology. This role directly impacts Crusoe's mission by enabling customers to accelerate their research and development, contributing to a more sustainable future. You will be involved in exciting projects, working with cutting-edge technologies and collaborating with a talented team to solve complex challenges. The ideal candidate is a highly motivated and experienced technical professional with a passion for customer success, a deep understanding of cloud technologies, and a commitment to Crusoe's values.

$145k - $175k

San Francisco, CA - US onsite FullTime
AWSAzureKubernetes +2 more

Sr. Cloud Support Engineer - Weekend Shift

1mo ago
c

crusoe

Crusoe Cloud is revolutionizing high-performance computing by offering sustainable, low-cost GPU compute power. As a Senior Cloud Support Engineer, you will play a crucial role in empowering our customers to leverage this technology for groundbreaking advancements in fields like AI/ML, physics simulations, and computational biology. You will be the primary point of contact for technical support, ensuring our customers can seamlessly utilize Crusoe Cloud to achieve their goals. This role directly impacts Crusoe's mission by enabling our customers to accelerate their research and development, contributing to a more sustainable future. You will be involved in exciting projects, working with cutting-edge technologies and collaborating with a talented team to solve complex challenges.

$135k - $165k

Denver, CO - US onsite FullTime
AWSAzureKubernetes +2 more

Research Engineer, Infrastructure, Inference

2mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking an infrastructure research engineer to design, optimize, and scale the systems that power large AI models. The goal is to make inference faster, more cost-effective, more reliable, and more reproducible, enabling research teams to focus on advancing model capabilities. This role is crucial for ensuring that every experiment, evaluation, and deployment runs smoothly at scale, with a focus on performant and efficient model inference for both real-world applications and research acceleration.

$350k - $475k

San Francisco onsite
OpenAIMistralKubernetes +3 more

Staff Software Engineer

2mo ago
f

fiddler-ai

Fiddler is building trust into AI, especially with the rise of Generative AI and Agents. Our platform helps organizations deploy trustworthy and transparent AI solutions by monitoring, evaluating, securing, analyzing, and improving AI applications. We partner with AI-first organizations to establish responsible AI practices, fostering user trust. Our Integrations Team focuses on the connective tissue between customer AI stacks and our observability platform, building OpenTelemetry-native SDKs, framework instrumentation, and ingestion pipelines for various AI environments including predictive models, LLMs, GenAI, and agentic applications. This role offers a broad scope, combining SDK/instrumentation leadership, OpenTelemetry expertise, distributed systems architecture, and AI observability platform development.

Bengaluru hybrid FullTime
LangGraphLangChainAWS +5 more

Member of Technical Staff - Engineering Lead, Platform Foundations

2mo ago
Reflection ai

Reflection ai

Reflection is seeking a Platform Foundations Lead to build and operate the foundational layer that all engineering teams depend on. This includes cloud and multi-cloud infrastructure, networking, security, and developer tooling. The role involves leading a team of infrastructure engineers, guiding technical and architectural decisions, and working closely with other engineering teams. You will also manage cloud and infrastructure vendors and contribute as an individual contributor to stay close to the technical stack.

San Francisco, CA onsite FullTime
KubernetesTerraform

Member of Engineering (Inference Infrastructure)

2mo ago
p

poolside

Poolside is building a world where AI drives economically valuable work and scientific progress, aiming to accelerate software development with agentic systems, coding assistants, and frontier models. This role is on the compute team, focusing on optimizing GPU workload scheduling and inference serving. You will partner with the inference team to enhance throughput and latency for evaluations and reinforcement learning, collaborate with the scalability team to stabilize large-scale fault-tolerant training, and work closely with the infrastructure team to ensure GPU nodes are healthy and fully utilized. Your work will directly impact research velocity and contribute to the company's mission of building frontier models.

Remote (EMEA) remote FullTime
KubernetesGoReinforcement Learning

Member of Engineering (Infrastructure)

2mo ago
p

poolside

Poolside is building a world where AI is the engine behind economically valuable work and scientific progress, aiming to accelerate software development with agentic systems, coding assistants, and frontier models. This role focuses on designing and implementing the build, CI/CD, and tooling infrastructure that researchers and engineers rely on daily. You will join the infrastructure team as the second engineer focused on developer experience, partnering closely with the current owner while sharing on-call, incident response, and platform ownership. The mission is to make it seamless for researchers and engineers to build, test, and ship by optimizing the monorepo, ensuring reproducible builds, and maintaining a coherent Python ecosystem through automation.

Remote (EMEA) remote FullTime
AWSKubernetesPython +2 more

Senior AI Solutions Engineer

2mo ago
f

fiddler-ai

Fiddler is building trust into AI, especially with the rise of Generative AI and Agents. Our platform helps organizations deploy trustworthy and transparent AI solutions by monitoring, evaluating, securing, analyzing, and improving AI applications. We partner with AI-first organizations to establish responsible AI practices, enabling engineering and business teams to understand AI outcomes. Joining Fiddler means making an impact by ensuring AI applications at production scale have operational transparency and security, offering monumental learning opportunities in a rapidly innovating industry.

$174k - $215000k

US - Remote remote FullTime
AWSKubernetesGo +5 more

Manager of Technology & Security Engineering

2mo ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible. We build open models that empower individuals to control their intelligence and shape the future of AI. The Head of Technology & Security Engineering will architect and operate the security engineering foundation protecting our corporate environment, multi-cloud research infrastructure, and significant GPU training capacity. This leadership role encompasses the full technical security stack, from end-user compute and zero-trust access to cloud security, detection engineering, and security-as-code, enabling rapid research progress without compromising rigor. Positioned at the intersection of security engineering, infrastructure, and research operations, this role is crucial for safeguarding sensitive assets like model weights, training data, and GPU capacity, while maintaining a low-friction environment for researchers and engineers. The ideal candidate possesses deep technical expertise and executive presence to lead the security engineering function, represent the company's security posture, and manage vendor and contractual risks.

New York, NY onsite FullTime
KubernetesPythonTerraform

Engineering Manager, Production Engineering

2mo ago
Harvey

Harvey

Harvey is seeking a Senior Engineering Manager to lead the Infrastructure Foundation & Production Quality Engineering organization. This team is crucial for building and operating Harvey's core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. The role involves owning the reliability, scalability, security, and efficiency of the infrastructure platform, leading a team of engineers, and partnering with various departments to ensure the infrastructure scales with rapid growth. This is a leadership role reporting to the Head of Infrastructure, focused on shaping the future of Harvey's infrastructure platform.

$260k - $340k

San Francisco hybrid FullTime
AWSAzureKubernetes +2 more

Member of Engineering (Agent Sandboxes)

2mo ago
p

poolside

Poolside is building a company to create Artificial General Intelligence, aiming to accelerate software development through agentic systems and frontier models. We are a distributed team across Europe and North America, with monthly in-person collaboration in Paris and annual off-sites. Our multidisciplinary team is united by a deep care for our work, fostering a culture of hard work, intellectual curiosity, and kindness. We are seeking a Member of Engineering to join our Platform engineering team, focusing on our sandboxed agent execution environment. This role is crucial for accelerating experimentation and foundation model training by designing, developing, and operating services and systems that enable running agent sessions safely and at scale.

Remote (EMEA/East Coast) remote FullTime
AWSKubernetesPython +2 more

Senior Software Engineer, Cloud Infrastructure

2mo ago
d

decagon

Decagon is seeking a Senior Software Engineer to join their Infrastructure team. This role will focus on building and operating the foundational systems that power Decagon's conversational AI platform, including networking, data, ML serving, and developer platforms. You will be responsible for architecting and operating deployments within enterprise customer clouds, ensuring reliability, security, and compliance. This position involves building platforms and abstractions for product teams, managing end-to-end ownership of enterprise deployments, and ensuring the reliability of agentic AI workloads. The role requires a proactive approach to problem-solving in a rapidly evolving technological landscape.

$200k - $400k

San Francisco onsite FullTime
AWSAzureKubernetes +3 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.