Kubernetes Jobs
410 open roles mentioning Kubernetes
Senior Systems Engineer, Workers AI
Cloudflare
You will design and build the core infrastructure that powers AI inference across Cloudflare's global network, handling real-time voice, frontier open LLMs, and customer-deployed models on a heterogeneous fleet of GPUs and next-generation accelerators. This role involves solving complex problems in distributed systems and high-performance computing, such as sub-second model cold starts, multi-accelerator workload scheduling, efficient KV cache management, and a model deployment platform. We are building a novel AI inference platform embedded in the internet's fabric and are seeking high-agency systems engineers who are passionate about foundational infrastructure and defining how AI operates at the network edge.
Software Engineer, Cloudflare Network Interconnect
Cloudflare
The CNI team builds and operates systems that enable Cloudflare's enterprise customers to connect their networks directly to Cloudflare via private, dedicated interconnects. This product, Cloudflare Network Interconnect (CNI), supports various interconnect types across hundreds of global locations and is foundational to key services like Magic Transit, Cloudflare WAN, Zero Trust, and application security for major clients. As a member of this team, you will develop and enhance the systems responsible for provisioning, monitoring, and orchestrating these private network interconnects at scale. You will collaborate with Product Management, Network Engineering, and Infrastructure teams on significant initiatives to integrate enterprise customers directly into Cloudflare's global network.
Software Engineer - Platforms & Productivity
Cloudflare
The Developer Productivity team at Cloudflare builds internal tools and platforms to enhance the velocity and developer experience for Cloudflare's engineering teams. This role is central to AI-assisted development across all of Cloudflare engineering, with significant adoption growth and a small team structure offering broad ownership. As a Software Engineer, you will contribute to building and operating internal platforms that enable Cloudflare engineers to develop, deploy, and run software safely and efficiently. Your work will span developer tooling, CI/CD, GitOps, and AI-assisted development, including MCP servers, agents, evals, and context engineering. You will also play a key role in governing and enforcing Cloudflare’s Engineering Codex by translating engineering standards into automated guardrails, paved paths, and actionable feedback.
Senior Data Engineer
Cloudflare
Cloudflare is seeking an experienced Data Engineer to join their Austin team. This role involves scaling the development of internal data products, building data-driven applications that empower various teams across the company, including go-to-market, engineering, and product. As the team initiates and owns its products, you will be involved end-to-end, from shaping requirements and designing features to implementation and long-term ownership. You will work on building scalable, reliable systems that solve critical business problems, partnering closely with full-stack engineers to develop new features and operate data pipelines and services. The role offers opportunities to work with a diverse stack including Go, Scala, and ClickHouse, and to build sophisticated, AI-ready data layers using tools like vector databases and Workers AI.
Solutions Architect, Customer Success - US (Remote)
fiddler-ai
Fiddler is seeking a Solutions Architect, Customer Success to ensure clients achieve significant, measurable outcomes from their AI observability investments. In this role, you will act as both a technical expert and a trusted advisor, connecting complex Machine Learning (ML) and Large Language Model (LLM) systems with tangible business value. By guiding customers through the onboarding process, building integrations, and advocating for their needs internally, you will help them deploy trustworthy AI at scale, contributing to Fiddler's growth through successful adoption, renewals, and expansion.
$160k - $200k
Senior Software Engineer, GPU Infrastructure (HPC)
Cohere
Cohere is seeking a Staff Software Engineer to join our internal infrastructure team, responsible for building and operating world-class infrastructure and tools for training, evaluating, and serving Cohere's foundational AI models. You will work closely with AI researchers to support their AI workload needs on cutting-edge systems, focusing on stability, scalability, and observability. This role involves building and operating superclusters across multiple clouds, directly accelerating the development of industry-leading AI models. Participation in a 24x7 on-call rotation is required and compensated.
Software Engineer, Developer Productivity
Glean
Glean is seeking a Software Engineer, Developer Productivity to enhance how our engineers build, test, and ship software. This role involves designing and optimizing build systems, CI/CD pipelines, and developer tooling within a large, multi-language codebase powered by Bazel and running on GitHub Actions with cloud remote execution. You will play a key part in reducing friction in daily workflows, scaling infrastructure, and enabling the entire engineering team to move faster and with greater confidence, including leveraging AI-powered coding and productivity tools.
$140k - $220k
Staff + Senior Software Engineer, Inference Infrastructure
Anthropic
Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide, managing the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team's dual mandate is to maximize compute efficiency for explosive customer growth while enabling breakthrough research by providing scientists with high-performance inference infrastructure for next-generation models. This involves tackling complex, distributed systems challenges across multiple accelerator families and emerging AI hardware on various cloud platforms.
Senior Platform Engineer
Legora
Legora is redefining how legal work gets done with an AI-native workspace designed for legal professionals. Our platform helps teams analyze documents rapidly, streamline workflows, and focus on strategic judgment and outcomes. Trusted by over 1,000 customers globally, including major law firms and corporations, Legora has achieved significant ARR growth and continues to expand through strategic acquisitions. We foster a culture of ownership, excellence, and continuous growth, where impact and pace are paramount. As a Senior Platform Engineer, you will join an enabling team dedicated to enhancing developer productivity by improving platform reliability, performance, and scalability. You will work hands-on with infrastructure and backend systems to empower teams to ship code faster, safer, and smarter.
Product Engineer, Cyber
OpenAI
We are seeking backend-focused Product Engineers to join our platform and security product teams. In this role, you will build infrastructure and customer-facing workflows that enable developers and AI agents to work reliably in parallel. You will own outcomes from understanding user problems and choosing approaches through shipping, operating, and improving solutions, working closely with frontend, infrastructure, and security engineers. This is an opportunity to build essential software factory components for enterprises, enabling AI agents to operate securely within customer-controlled cloud environments.
Staff+ Software Engineer, Infrastructure, Interpretability
Anthropic
Anthropic is building reliable, interpretable, and steerable AI systems to ensure AI is safe and beneficial for society. The Interpretability team focuses on understanding the inner workings of trained models and applying techniques to enhance AI safety. This role is an early hire on a new infrastructure effort within the Interpretability team, focused on defining and building the systems that enable secure, private, and low-friction access to frontier models for researchers. The work involves designing and implementing solutions across security, privacy, data and compute management, and developer experience to support cutting-edge AI research and its application to safety decisions.
Software Engineer - Identity & Authorization
Baseten
Baseten powers mission-critical inference for leading AI companies, enabling them to bring cutting-edge models into production. We are seeking a founding engineer for our identity and authorization team within enterprise engineering. This role will own the identity and access layer of the Baseten platform, including the authorization model, credential systems, and admin experiences for enterprise IT teams. You will design and build a fine-grained authorization system from the ground up, ensuring low-latency permission checks at high request volumes, consistent behavior across the product suite, and strong security guarantees for critical workloads.
Engineering Manager, FDE Infrastructure (NORAM)
Cohere
Cohere is seeking an Engineering Manager to lead their Deployment Engineering team. This is a hands-on leadership role focused on driving the deployment of Cohere's North platform into customer environments. The manager will lead a team of Forward Deployed Engineers, taking ownership of technical implementation and customer success. The role involves close collaboration with Product, Engineering, and Sales teams to deliver AI solutions to enterprises and requires a proactive approach to problem-solving and scaling infrastructure.
Forward Deployed Engineer, Infrastructure Specialist (Public Sector)
Cohere
Cohere is seeking a Forward Deployed Engineer, Infrastructure Specialist to partner with Canadian public sector organizations. This role involves the secure and ethical deployment of Generative AI solutions, working directly with clients to understand their challenges and implement solutions using Cohere's platform. You will act as a bridge between Cohere's core product and client engineering teams, focusing on solving complex problems and securely integrating AI into critical sectors. The ideal candidate will have diverse skill sets in backend, infrastructure, agent development, and deployments, with a strong customer focus and a passion for working at the cutting edge of Agentic AI.
$20k - $40k
Forward Deployed Engineer, Infrastructure Specialist (North America)
Cohere
Cohere is a leading enterprise AI company building cutting-edge foundation AI models and end-to-end products. We are seeking engineers to join our team and contribute to the widespread adoption of AI. This role offers a unique opportunity to shape how enterprises harness the power of AI in real-world applications, acting as a bridge between our core North product and client engineering teams. You will be at the forefront of solving complex problems and securely integrating AI into critical sectors.
GRC Program Manager, Product Lifecycle Assurance
OpenAI
OpenAI is seeking an experienced Product Lifecycle Assurance Individual Contributor to scale its Governance, Risk, and Compliance (GRC) function across its product stack. This role is crucial for ensuring products meet customer and regulatory compliance requirements at launch, and for promptly detecting and correcting any regressions. You will partner closely with Product, Security, Legal, and Privacy teams to maintain OpenAI's security, privacy, and compliance claims, providing assurance to customers, auditors, and regulators regarding user data handling. The position involves end-to-end product assurance, from inception through continuous post-launch monitoring, enhancing and building components for a comprehensive program. This is a highly cross-functional and technical operations role focused on ensuring products meet compliance standards at launch, preventing regressions, and evidencing compliance state, rather than supporting traditional audits.
Technical Account Manager
Lambda
Lambda, a leader in AI cloud infrastructure, is seeking a Technical Account Manager to own the technical health of post-sales relationships for their public cloud accounts. This hands-on role involves understanding customer AI use cases, leading joint POC sessions, designing and defending architectures, and owning SLA and reliability engineering. You will build tooling to make accounts measurable, direct technical escalations, and serve as the technical voice of the customer within Lambda. The ideal candidate will have a deep understanding of GPU infrastructure and AI/ML workloads, with a proven ability to lead structured technical engagements and communicate complex technical content effectively.
Security Engineer - Threat Intel
Anthropic
Anthropic is at the forefront of AI development, making it a prime target for sophisticated adversaries. The Threat Intelligence function within the Detection & Response team is crucial for anticipating and mitigating these threats. As a Threat Intelligence Engineer, you will be a hands-on practitioner responsible for generating actionable intelligence that guides our detection strategies, threat hunts, and defensive priorities. You will track adversaries targeting frontier AI labs, develop tools and pipelines to transform raw indicators into operational defenses, and collaborate closely with detection engineers and incident responders to ensure intelligence effectively influences security outcomes. This is a high-impact, builder-oriented role within a small team, offering significant autonomy to shape the collection, analysis, and operationalization of threat intelligence at Anthropic.
Staff + Sr. Software Engineer, Scaling
Anthropic
Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. The Inference team is responsible for building and scaling the critical systems that serve Claude to millions of users worldwide. This involves managing the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators, with a dual mandate of maximizing compute efficiency for customer growth and enabling breakthrough research by providing high-performance inference infrastructure. The team tackles complex, distributed systems challenges across multiple accelerator families and emerging AI hardware on various cloud platforms.
Engineering Manager, Inference Infrastructure
Anthropic
Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. This role leads a team of ML platform, infrastructure, and distributed-systems engineers responsible for the critical control plane that manages Anthropic's inference fleet. This involves making key decisions about request routing, capacity allocation, and system performance to meet throughput, reliability, and latency constraints. The team designs placement and load-balancing algorithms, builds quantitative models for demand and capacity, and optimizes latency across various system boundaries. The Engineering Manager will own the technical roadmap, drive quantitative modeling habits, set technical strategy for evolving the control plane, and ensure the operational health of the inference request path.