Kubernetes Jobs
204 open roles mentioning Kubernetes
Senior Software Engineer, Production Engineering
Harvey
Harvey is seeking a Production Engineer to help build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth. The ideal candidate will have a systems-thinking mindset and a passion for building simple, reliable, and scalable systems.
$161k - $242k
Staff Software Engineer, Production Engineering
Harvey
Harvey is seeking a Production Engineer to build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth.
$231k - $340k
Staff Software Engineer, Production Engineering
Harvey
Harvey is seeking a Production Engineer to build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth.
$231k - $340k
Senior Software Engineer, Production Engineering
Harvey
Harvey is seeking a Production Engineer to help build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth. The ideal candidate will have a systems-thinking mindset and a passion for building simple, reliable, and scalable systems.
$161k - $242k
Infrastructure engineer (UK)
Writer
At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need.
Infrastructure engineer
Writer
At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need.
Engineering Manager, Model Infrastructure
Harvey
Harvey is transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. This role offers a rare chance to help build a generational company at a true inflection point, with strong product-market fit and world-class investor support. The team moves fast, takes ownership, and is deeply committed to the mission, operating with intensity and pushing for excellence. As the Engineering Manager for Model Infrastructure, you will lead the team responsible for the platform powering every model request across Harvey, partnering closely with AI Research, Product Engineering, and Infrastructure to ensure reliability, scalability, and cost-efficiency. This is a strategic engineering organization, critical for every product capability, and will evolve to build the infrastructure for Harvey to train, evaluate, deploy, and operate its own frontier AI models.
$260k - $340k
Engineering Manager, GPU Infrastructure
Cohere
Cohere is a leading enterprise AI company building cutting-edge foundation AI models and end-to-end products. The GPU Clusters team is central to Cohere's infrastructure, responsible for building and operating the superclusters that power our frontier AI models. This role involves enabling research and development at the intersection of cutting-edge hardware, distributed systems, and AI research. As an Engineering Manager, you will lead a team of highly motivated engineers passionate about GPU infrastructure and AI, fostering a culture of technical excellence and innovation in a remote-first environment. This is a unique opportunity to shape the infrastructure powering the next generation of AI.
Staff+ Software Engineer, Platform
Anthropic
Anthropic is seeking experienced software engineers to join its Platform organization. This team builds the foundational primitives that accelerate product development and owns the infrastructure and systems that enable reliable, scaled product shipping for internal and external users. You will independently scope complex, multi-month projects, drive cross-organizational alignment, and make architectural decisions that shape how Anthropic builds and scales its products. You will partner directly with research to productize cutting-edge capabilities, impacting the platform used by hundreds of thousands of companies and engineers daily.
Software Engineer, Developer Productivity, AI Tools
thinkingmachines
Thinking Machines Lab is seeking a developer productivity engineer to enhance internal software development processes, focusing on safety, speed, and user experience. This role will concentrate on AI tools and coding agents, collaborating with platform, security, and product engineers to build cutting-edge tooling for AI-assisted software development and significantly accelerate the inner development loop. The position involves both establishing company-wide platforms and assisting individual developers in optimizing their workflows.
$350k - $475k
Site Reliability Engineer (SRE)
thinkingmachines
Thinking Machines Lab is seeking a Site Reliability Engineer (SRE) to ensure the end-to-end reliability of their Tinker platform. This role involves working closely with engineers and research teams to enhance the robustness and resilience of every system layer. The SRE will be instrumental in maintaining and improving the infrastructure that supports custom AI model fine-tuning, ensuring a seamless experience for researchers and developers.
$350k - $475k
Staff+ Software Engineer, Capacity Engineering
Anthropic
Anthropic is building reliable, interpretable, and steerable AI systems to ensure AI is safe and beneficial. The Capacity Engineering team plays a crucial role in managing one of the industry's largest and fastest-growing infrastructure fleets. This team is responsible for ensuring all infrastructure resources are accounted for, well-utilized, and efficiently allocated. As an engineer on this team, you will develop the production systems that power these efforts, including data pipelines for telemetry ingestion, observability tools for real-time fleet health visibility, and performance instrumentation to measure hardware utilization across various workloads. You will write production-quality code, operate at scale with Kubernetes-native infrastructure, and significantly influence decisions regarding Anthropic's substantial infrastructure spend. This role involves close collaboration with research engineering, infrastructure, inference, and finance teams, requiring comfort across data engineering, systems engineering, and observability in a high-autonomy environment.
Research Engineer, Infrastructure, Inference
thinkingmachines
Thinking Machines Lab is seeking an infrastructure research engineer to design, optimize, and scale the systems that power large AI models. The goal is to make inference faster, more cost-effective, more reliable, and more reproducible, enabling research teams to focus on advancing model capabilities. This role is crucial for ensuring that every experiment, evaluation, and deployment runs smoothly at scale, with a focus on performant and efficient model inference for both real-world applications and research acceleration.
$350k - $475k
Senior Software Engineer, Inference
Anthropic
Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry's largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team has a dual mandate: maximizing compute efficiency to serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms.
Staff Software Engineer, Inference
Anthropic
Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry's largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team has a dual mandate: maximizing compute efficiency to serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms. As a Staff Software Engineer on our Inference team, you will work end to end, identifying and addressing key infrastructure blockers to serve Claude to millions of users while enabling breakthrough AI research.
Member of Technical Staff - Engineering Lead, Platform Foundations
Reflection ai
Reflection is seeking a Platform Foundations Lead to build and operate the foundational layer that all engineering teams depend on. This includes cloud and multi-cloud infrastructure, networking, security, and developer tooling. The role involves leading a team of infrastructure engineers, guiding technical and architectural decisions, and working closely with other engineering teams. You will also manage cloud and infrastructure vendors and contribute as an individual contributor to stay close to the technical stack.
Manager of Technology & Security Engineering
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible. We build open models that empower individuals to control their intelligence and shape the future of AI. The Head of Technology & Security Engineering will architect and operate the security engineering foundation protecting our corporate environment, multi-cloud research infrastructure, and significant GPU training capacity. This leadership role encompasses the full technical security stack, from end-user compute and zero-trust access to cloud security, detection engineering, and security-as-code, enabling rapid research progress without compromising rigor. Positioned at the intersection of security engineering, infrastructure, and research operations, this role is crucial for safeguarding sensitive assets like model weights, training data, and GPU capacity, while maintaining a low-friction environment for researchers and engineers. The ideal candidate possesses deep technical expertise and executive presence to lead the security engineering function, represent the company's security posture, and manage vendor and contractual risks.
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering
Anthropic
Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. The Cloud Inference team is responsible for scaling and optimizing Claude to serve massive audiences across AWS, GCP, Azure, and other cloud providers. This involves end-to-end ownership of Claude on each platform, from API integration and request routing to inference execution and capacity management. The model & inference launch team specifically focuses on the validation pipeline for the inference server and load balancer, ensuring every inference change, including model launches and performance improvements, is deployed with correctness, performance, and reliability. This high-leverage infrastructure work is critical for fast and cost-effective deployment of frontier models and features, directly impacting compute capacity and the speed at which innovations reach production.
Staff Software Engineer, Inference
Anthropic
Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry's largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team has a dual mandate: maximizing compute efficiency to serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms.
Technical Program Manager, Infrastructure
Anthropic
Anthropic's Infrastructure organization is the engine that powers our mission, building and operating massive clusters for training frontier models, production infrastructure serving millions of users reliably, and developer platforms. As a Technical Program Manager for Infrastructure, you will coordinate complex programs with broad organizational impact, solving novel scaling challenges while maintaining security and reliability. This role is ideal for someone who thrives in ambiguity, makes others more effective, and partners closely with engineering leadership to drive strategic initiatives and ensure seamless coordination between research, engineering, and product teams.