Kubernetes Jobs
410 open roles mentioning Kubernetes
Senior Backend Software Engineer - Core Backend, Cloud Customer Experience
crusoe
Crusoe is seeking a Senior Backend Software Engineer to join the Core Backend team within Cloud Customer Experience (CCX). This role is crucial for delivering a best-in-class user experience for Crusoe's AI-focused cloud platform. The Core Backend team is responsible for the foundational services that power every customer interaction, ensuring reliability and scalability. The work involves building and operating key systems like the API Gateway, Resource Management Layer, Quota Management, Notifications, and User Onboarding Flows, all while focusing on speed, reliability, and coherence as the platform grows.
$170k - $205k
Application Security Engineer
Glean
Glean is seeking an experienced Application Security Engineer to ensure the security of its technology stack by eliminating software vulnerabilities. This role will be responsible for securing base OS images, scanning and patching open-source software (OSS) dependencies, and integrating advanced security tools into the CI/CD pipeline. The ideal candidate will lead Glean's vulnerability management efforts, driving the adoption of solutions like Google's Assured Open Source Software (OSS) and exploring new technologies and processes to proactively protect the infrastructure.
$185k - $260k
Platform Security Engineer
Glean
Glean is seeking a talented security-focused software engineer to join our growing team. In this role, you will play a critical part in developing and maintaining the security foundation of our platform. You will be responsible for designing, implementing, and testing security features across various software components, ensuring the integrity and safety of our Work AI platform.
$185k - $260k
Supply Chain Security Engineer
Glean
Glean is seeking a Supply Chain Security Engineer to ensure the security of our technology stack by eliminating software vulnerabilities (CVEs). This role involves securing base OS images, scanning and patching open-source software (OSS) dependencies, and integrating advanced security tools into our CI/CD pipeline. You will be instrumental in building and executing our software supply chain security strategy, enhancing vulnerability scoring, and reducing the overall vulnerability footprint across various programming ecosystems and base layers. The position also includes creating hardened images, leading initiatives like SBOM generation and automated fix pipelines, and developing policy-driven controls for build provenance and trust verification. Additionally, you will manage secure software supply chain guidelines and documentation, and contribute to achieving FEDRAMP readiness for vulnerability management.
Director, Product Management
Runpod
Runpod is seeking a Director, Product Management to define and execute the product vision for its AI cloud platform. This role will oversee the entire product lifecycle for core offerings like serverless endpoints, GPU instances, AI foundational services, developer APIs, and the overall user experience. The successful candidate will build and lead a high-performing product and design team, deeply understanding the needs of AI researchers, ML engineers, and developers. This position requires close collaboration with Infrastructure, Engineering, and Go-to-Market leadership to ensure Runpod remains the fastest, most intuitive, and cost-effective platform for training and running large-scale AI inference workloads.
$225k - $325k
Staff + Senior Software Engineer, Inference
Anthropic
The Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. This role involves managing the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team's dual mandate is to maximize compute efficiency for explosive customer growth while enabling breakthrough research by providing high-performance inference infrastructure for next-generation models. This involves tackling complex, distributed systems challenges across multiple accelerator families and emerging AI hardware on various cloud platforms.
Research Engineer, Pretraining
Anthropic
Anthropic is seeking a Research Engineer to join its Pretraining team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, aiming to build safe, steerable, and trustworthy AI systems. The mission is to ensure that transformative AI systems are aligned with human interests and are beneficial for society. The team is dedicated to pushing the boundaries of AI while prioritizing safety and ethics.
[Expression of Interest] Research Engineer / Scientist, Alignment - London
Anthropic
Anthropic is building reliable, interpretable, and steerable AI systems to ensure AI is safe and beneficial for society. As a Research Engineer on the Alignment Science team in London, you will design and execute machine learning experiments to understand and steer the behavior of advanced AI systems. You will focus on AI safety, particularly risks from future human-level AI systems, collaborating with teams like Interpretability and Frontier Red Team. The role involves exploratory research in areas such as AI Control and Alignment Stress-testing, aiming to make AI helpful, honest, and harmless.
Staff Software Engineer, Inference
Anthropic
Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry's largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team has a dual mandate: maximizing compute efficiency to serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms. As a Staff Software Engineer on our Inference team, you will work end to end, identifying and addressing key infrastructure blockers to serve Claude to millions of users while enabling breakthrough AI research.
Staff Software Engineer, Inference
Anthropic
Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry's largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team has a dual mandate: maximizing compute efficiency to serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms.
Deployed Architect, Professional Services (APAC)
Langchain
LangChain is seeking a Deployed Architect to join our Professional Services team. In this role, you will collaborate directly with enterprise clients to design, implement, and refine production-grade AI infrastructure and agent systems. Your responsibilities will include architecting robust, secure infrastructure deployments and developing reliable, well-evaluated agent applications tailored to address specific business challenges. This position uniquely blends software development, infrastructure/platform engineering, and client-facing expertise, requiring deep technical knowledge in both infrastructure and agent engineering.
Deployed Architect, Professional Services (Amsterdam)
Langchain
LangChain is seeking a Deployed Architect to join our Professional Services team. In this role, you will collaborate directly with enterprise clients to design, implement, and fine-tune production-ready AI infrastructure and agent systems. Your responsibilities will include architecting scalable and secure infrastructure deployments, as well as building reliable and well-evaluated agent applications to address significant business challenges. This position uniquely blends software development, infrastructure/platform engineering, and client-facing expertise, requiring deep technical knowledge in both infrastructure and agent engineering.
Deployed Architect, Professional Services (London)
Langchain
LangChain is seeking a Deployed Architect to join our Professional Services team. In this role, you will collaborate directly with enterprise clients to design, implement, and fine-tune production-ready AI infrastructure and agent systems. Your responsibilities will include architecting robust, secure infrastructure deployments and developing reliable, well-evaluated agent applications to address critical business challenges. This position blends software development, infrastructure/platform engineering, and client interaction, requiring deep technical expertise in both infrastructure and agent engineering.
Member of Technical Staff (Software Engineer, API Platform)
Perplexity AI
Perplexity is seeking strong engineers with a passion for delivering frontier intelligence and infrastructure to developers. Our company builds technology that reshapes how people search, reason, and interact with the world around them. The API Platform engineering team is charged with designing, implementing, and scaling programmatic interfaces to that technology. As a member of our team, you'll work on an eclectic portfolio spanning distributed systems, performance optimization, agent orchestration, and frontier topics that often change with each passing month. Throughout this work, you'll prioritize great developer and agent experience alike, and define technical strategy for how we scale to meet compounding growth exponentials.
Software Engineer, Agents
Mercor
Mercor is seeking a strong engineer to build scalable agentic products. You will work on a variety of teams focused on automating operational work, developing evaluation systems for RL environments, and building the infrastructure for producing frontier data. This role involves end-to-end ownership of agentic features, from initial scoping and design through implementation, launch, and iteration. You will collaborate with researchers, operators, and AI companies at the forefront of AI advancement, contributing to the development of systems that are redefining society.
AI Engineer, Enablement
Langchain
LangChain is seeking an AI Engineer, Enablement to set the technical foundation for how customers learn to build reliable agents with the LangChain ecosystem. This role involves teaching teams to work effectively with LangChain, LangGraph, Deep Agents, and LangSmith through instructor-led workshops, written content, and reference implementations. You will collaborate closely with the broader Go-To-Market organization to ensure customers have the skills and confidence to build independently. The ideal candidate has experience building real agent systems, understands their tradeoffs, and possesses a passion for teaching, whether through large workshops or one-on-one debugging sessions. You will also contribute to building internal agents and tools to enhance the Enablement team's efficiency.
$150k - $195k
Senior Cloud Support Engineer
crusoe
Crusoe Cloud is seeking a Senior Cloud Support Engineer to empower customers utilizing their sustainable, low-cost GPU compute power for AI/ML, physics simulations, and computational biology. In this role, you will be the primary technical support contact, ensuring customers can seamlessly leverage Crusoe Cloud. This position is crucial for enabling customers to accelerate their research and development, contributing to a more sustainable future and Crusoe's mission. You will work with cutting-edge technologies and collaborate with a talented team to solve complex challenges, making it an impactful role for a motivated and experienced technical professional passionate about customer success and cloud technologies.
$145k - $175k
Customer Success Manager, Managed Inference
crusoe
Crusoe is seeking a motivated Customer Success Manager to support customers running AI inference workloads on their cloud platform. This role, within the Customer Experience organization, will help clients navigate the operational and technical challenges of deploying and scaling AI applications, from model serving and GPU utilization to production readiness. You will foster strong customer relationships, act as a liaison to technical teams, and assist clients in maximizing the value derived from Crusoe's AI and ML solutions. This is a full-time position based in the Bay Area, CA.
$175k - $200k
Software Engineer, Marketplace
Mercor
As a Software Engineer on the Marketplace team, you will own the core systems that bring human intelligence to AI opportunities. You'll work on search, matching, allocation, and workflow infrastructure at the heart of the marketplace, where systems deliver the world's top experts to staff cutting-edge projects. The problems are technically demanding and tightly coupled to business outcomes: latency, reliability, throughput, and system quality all directly affect marketplace performance. This role sits on a core decision layer of the product, and your work will directly shape how the marketplace operates, including which opportunities get filled, how quickly hiring happens, and how reliably the system scales as volume and complexity increase.
Member of Technical Staff (Software Engineer, Model Platform)
Perplexity AI
Perplexity AI is seeking a deeply technical software engineer to own and evolve its model serving platform. This mission-critical system connects products and research systems with inference layers, providing a fast and reliable interface across providers. The role involves working at the intersection of distributed systems, AI products, and infrastructure, making it easier for teams to adopt new models, run experiments, and ship dependable products. The ideal candidate will possess strong technical judgment, curiosity about frontier models, and experience designing core abstractions and leading technical decisions.