Kubernetes Jobs
204 open roles mentioning Kubernetes
Staff + Senior Software Engineer, Inference Deployment
Anthropic
Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. Our growing team of researchers, engineers, and policy experts collaborates to create these advanced AI systems. The Inference team is central to this mission, responsible for building and maintaining the critical systems that serve Claude to millions of users globally. We manage the entire stack, from intelligent request routing to fleet-wide orchestration across diverse AI accelerators, ensuring Claude is brought to life efficiently and reliably.
Engineering Manager, Production Engineering
Harvey
Harvey is seeking a Senior Engineering Manager to lead the Infrastructure Foundation & Production Quality Engineering organization. This team is crucial for building and operating Harvey's core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. The role involves owning the reliability, scalability, security, and efficiency of the infrastructure platform, leading a team of engineers, and partnering with various departments to ensure the infrastructure scales with rapid growth. This is a leadership role reporting to the Head of Infrastructure, focused on shaping the future of Harvey's infrastructure platform.
$260k - $340k
Security Operations Lead
Replit
Replit is seeking a Security Operations Lead (SOC Lead) to establish, enhance, and manage a 24/7 detection and response capability within a modern, cloud-native, and AI-driven environment. This leadership role will oversee global SOC functions including monitoring, SIEM management, detection engineering, alert triage, and operational readiness. A key aspect of this position involves evaluating and integrating emerging AI-based SOC products and autonomous response platforms. The role requires monitoring across multi-cloud environments (primarily GCP, with AWS/Azure secondary), Kubernetes, SaaS services, endpoints, developer tools, and AI workloads. Collaboration with Cloud Security, Compliance/GRC, SRE, Platform Engineering, IT/Endpoint teams, and AI Infrastructure is essential to ensure the detection strategy scales effectively against evolving threats. This is a hands-on leadership opportunity to shape the future of SOC operations while addressing complex challenges in a high-scale AI setting.
Senior/Staff Infrastructure Security Engineer
Abridge
Abridge is seeking a highly experienced and motivated Senior or Staff Infrastructure Security Engineer to join their founding security team. This role offers a unique opportunity to build security from the ground up at the forefront of AI in healthcare. You will be a key technical leader, shaping the company's product, infrastructure, and engineering practices by integrating security seamlessly, automating capabilities, and designing high-scale data pipelines. The position requires deep technical expertise, a builder's mindset, and excellent communication skills to foster a secure-by-default culture. This is a greenfield opportunity to architect the future of security at Abridge, focusing on large-scale data and automation challenges.
Machine Learning Engineer, Reliability
Fal
fal is building the generative media ecosystem for the next generation of AI products, providing the infrastructure, tools, and model access needed to scale from idea to production. As generative media reshapes industries, fal is becoming the foundation for ambitious teams. This hybrid ML Engineering / Site Reliability Engineering role will own the reliability, security, and safety of fal's generative media model APIs, ensuring they remain available, performant, secure, and safe for thousands of developers and enterprises. You will address model-specific failure modes, such as degraded output quality, drift, unsafe generations, and abuse patterns, as critical reliability concerns alongside uptime and latency.
Senior Software Engineer - Agentic Tooling & Productivity
Scale AI
Scale AI is seeking a Senior Software Engineer to design, build, and operate secure, scalable infrastructure that empowers employees. You will join a team focused on automating identity and access management, endpoint management, and the broader SaaS stack. The role involves leveraging internal and external models to create applications, slackbots, and dashboards for internal users, streamlining complex workflows through automation, and significantly enhancing the velocity of G&A teams. As a full-stack engineer, you will contribute to building internal and external tools, developing cloud-native distributed systems, integrating with SaaS platforms, and ensuring product quality through testing and debugging.
$216k - $270k
Deployment Engineering Manager, Enterprise
Scale AI
Scale AI is seeking an Infrastructure Engineering Manager to lead a team focused on the full deployment lifecycle of AI applications in live customer environments. This role bridges the gap between AI research and production, transforming innovative prototypes into scalable, high-performance enterprise solutions. The team works on interactive AI applications, enterprise SaaS products, and platform capabilities, ensuring deployments are production-ready, secure, and built to last. The manager will drive improvements through customer feedback, automation, and scalable playbooks, focusing on complex technical problems at the intersection of customer impact and platform excellence.
$216k - $270k
Senior Manager, Engineering
Harvey
Harvey is seeking an Engineering Manager to lead its Core Infrastructure team in Bengaluru. This role is crucial for scaling the platform, team, and future Harvey products, focusing on building new systems and hardening existing ones with an emphasis on architecture, scale, and long-term platform leverage. The Core Infrastructure team is responsible for delivering the foundational systems that power all user interactions across Harvey's legal AI platform, processing billions of prompt tokens and millions of daily requests. This is a high-leverage position for a builder-leader who can grow engineers, drive technical execution, and enhance reliability, performance, security, and developer velocity.
Forward Deployed Engineer, Infrastructure Specialist
Cohere
Cohere is seeking a Forward Deployed Engineer, Infrastructure Specialist to join their team. This role focuses on shaping how enterprises harness AI in real-world applications, acting as a bridge between Cohere's North AI workspace platform and client engineering teams. You will be at the forefront of solving complex problems and securely integrating AI into critical sectors like finance, healthcare, and telecommunications. The ideal candidate will be passionate about working at the cutting edge of Agentic AI and deeply care about customers.
$20k - $40k
Software Engineer, Adoption
Cohere
Cohere is seeking a Software Engineer to join the Agents & Automations team, which builds the core platform for North, an AI workspace designed for enterprises. This role involves developing AI-powered workflows, structured automations, and flexible agents that interact with sensitive enterprise data. You will contribute to building the workflow builder, execution engine, integrations, debugging tools, observability, evaluation systems, and feedback loops that enable customers to understand and improve their AI deployments. The work is broad, spanning frontend, backend, and AI-powered systems, with a focus on shipping reliable software that helps customers augment or automate business workflows with trustworthy agents and automations.
Senior/Staff Software Engineer, Developer Experience
Cohere
Cohere is seeking a Senior/Staff Engineer to build and maintain the automation infrastructure for their North platform. This role involves designing and implementing robust automation systems to enable efficient testing and validation across diverse environments. You will build systems, frameworks, and foster a culture where engineering teams can own quality themselves, improving the testing platform and empowering teams to ship with greater confidence. This is a critical role for ensuring the reliability and scalability of the North platform, offering the opportunity to shape the future of development infrastructure and significantly impact product quality.
Software Engineer, AI Infra
pika
We are seeking a Staff/Lead Software Engineer, AI Infrastructure, to play a critical role in building and scaling the core infrastructure that powers Pika’s AI capabilities. In this position, you will lead the design and implementation of GPU infrastructure, AI model serving APIs, and general AI infrastructure execution, enabling cutting-edge machine learning features that drive our products. You will be responsible for architecting robust, distributed systems optimized for high-performance AI workloads, large-scale GPU orchestration, and low-latency, reliable API serving. Your work will directly impact the way users experience and interact with generative AI at scale. As a senior technical leader, you’ll also mentor engineers, drive best practices, and set the technical vision for AI infrastructure at Pika.
AI Infrastructure Engineer
Together AI
As an AI Infrastructure Engineer at Together, you will be responsible for ensuring the smooth operation of all user-facing services and production systems. This role blends pragmatic operations with software engineering, applying sound engineering principles, operational discipline, and mature automation to our operating environments and codebase. You will specialize in systems (operating systems, storage subsystems, networking), implementing best practices for availability, reliability, and scalability, with varied interests in algorithms and distributed systems. Together AI is a research-driven artificial intelligence company focused on lowering the cost of modern AI systems through co-designing software, hardware, algorithms, and models, and we invite you to join our passionate group of researchers and engineers in building the next generation of AI infrastructure.
$190k - $270k
AI infrastructure Engineer (SRE) Amsterdam
Together AI
Together AI is seeking an AI Infrastructure Engineer (SRE) to ensure the smooth operation of user-facing services and production systems. This role combines the skills of a pragmatic operator and a software engineer, applying sound engineering principles, operational discipline, and automation to our operating environments and codebase. You will specialize in systems such as operating systems, storage subsystems, and networking, while implementing best practices for availability, reliability, and scalability, with interests in algorithms and distributed systems. Join a research-driven artificial intelligence company focused on advancing AI through open and transparent systems, aiming to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models.
LLM Inference Frameworks and Optimization Engineer
Together AI
Together.ai is building state-of-the-art infrastructure for efficient and scalable inference of large language models (LLMs). The company's mission is to optimize inference frameworks, algorithms, and infrastructure to push the boundaries of performance, scalability, and cost-efficiency. They are seeking an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines for multimodal and language models at scale. This role will focus on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design, ensuring efficient large-scale deployment of LLMs and vision models. This position offers a unique opportunity to shape the future of LLM inference infrastructure and ensure scalable, high-performance AI deployment across diverse applications.
$160k - $230k
Machine Learning, Platform Engineer
Together AI
Together AI is a research-driven artificial intelligence company focused on lowering the cost of modern AI systems. This role is part of a team dedicated to enabling custom models and dedicated inference on Together's platform. The team is responsible for building a container platform, optimizing autoscaling, minimizing cold starts, achieving the best end-to-end model performance, and providing a best-in-class developer experience with great tooling. The work often involves video or audio generation across the stack, including CUDA kernels, PyTorch optimization, inference engines, container orchestration, and queueing theory.
$160k - $250k
Research Engineer, Frontier Speculative Decoding
Together AI
Together AI is building the Inference Platform that powers the world's most advanced generative AI models. This role will serve as a critical bridge between cutting-edge research and real-world applications, focusing on translating internal model training research into production-ready deployments for customers. The work involves a deep commitment to data-centric development, meticulous hyperparameter tuning, and rigorous checkpoint evaluation. You will transform general-purpose models into highly performant, specialized tools by fine-tuning them on customer-specific data and internal datasets, working with dedicated GPU clusters rather than training foundation models from scratch.
$190k - $270k
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Together AI
Together AI is seeking a Staff Engineer to design and deliver multi-petabyte storage systems optimized for large-scale AI training and inference workloads. You will architect high-performance parallel filesystems and object stores, integrate cutting-edge technologies, and drive significant cost optimization. The role involves building Kubernetes-native storage operators and self-service platforms for automated provisioning and multi-tenancy. You will focus on optimizing data paths, designing multi-tier caching architectures, and tuning parallel filesystems for AI applications. This is a research-driven role within a company focused on lowering the cost of modern AI systems through co-design of software, hardware, algorithms, and models.
$250k - $300k
Senior Platform Engineer, Voice AI
Together AI
Together AI is building the best inference infrastructure for voice applications, powering production-grade, real-time voice agents and applications with best-in-class latency and reliability. We are seeking a Senior Platform Engineer to take ownership of the API and infrastructure layer for voice workloads. You will develop the real-time WebSocket and HTTP APIs used by developers to deploy voice experiences, design autoscaling for latency-sensitive streaming workloads, and ensure the reliability of our multi-provider voice platform for production voice agents handling millions of calls. This is a critical, foundational role on a small, high-impact team, defining how developers interact with our voice platform as we scale.
$200k - $260k
Senior Backend Engineer, Inference Platform
Together AI
Together AI is building the Inference Platform to bring advanced generative AI models to the world, powering multi-tenant serverless workloads and dedicated endpoints. This role offers a unique opportunity to optimize latency and fully utilize tens of thousands of GPUs, working hands-on with cutting-edge hardware. You will collaborate directly with research teams to productionize frontier models and engage with the open-source community, contributing to projects that push the boundaries of inference performance and efficiency.
$160k - $250k