Kubernetes Jobs

410 open roles mentioning Kubernetes

Security Engineer - Vuln Management (Infra)

3mo ago
Replit

Replit

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are seeking a mid-level Infrastructure Vulnerability Management Engineer with a strong background in Cloud Security, DevSecOps, and Infrastructure-as-Code (IaC). In this role, you will bridge the gap between security, compliance, DevOps, and Platform engineering teams. You will identify infrastructure misconfigurations, secure multi-cloud environments, and manage continuous vulnerability lifecycles across cloud workloads, containers, and data repositories to satisfy strict regulatory compliance frameworks. You will also serve as a technical infrastructure responder during security incidents, deploying real-time cloud or network countermeasures to protect our production ecosystem. What You'll Do Core Responsibilities - Infrastructure Scanning & Triage: Perform continuous security scanning across our cloud posture and workloads. Review, validate, and prioritize flaws and misconfigurations based on CVSS scores, real-world exploitability, and infrastructure network exposure. - Posture Management & Visibility: Own and optimize Cloud Security Posture Management (CSPM), Kubernetes Security Posture Management (KSPM), and Data Security Posture Management (DSPM) tools to ensure uniform compliance, prevent data leakage, and maintain hardened baselines. - Infrastructure-as-Code (IaC) Security: Configure, tune, and embed automated IaC security scanning tools into CI/CD pipelines to identify architectural risks (e.g., overly permissive IAM, public S3 buckets/Cloud Storage) before they are deployed to production. - Workload & Container Security: Manage the continuous vulnerability scanning lifecycle for container images, registries, and Virtual Machines (VMs), partnering with SRE and Platform teams to build automated base-image patching and rolling upgrade pipelines. - Compliance-Driven Tracking: Track, document, and manage infrastructure vulnerabilities according to strict compliance SLAs (e.g., SOC 2, ISO 27001, PCI-DSS). Maintain audit-ready evidence of infrastructure remediation timelines and exception approvals. - Executive Reporting & Alerting: Escalate and report critical production exposures directly to the CISO and senior leadership. Maintain dashboards and alerting mechanisms that visualize infrastructure risk trends and cloud compliance posture. - Remediation Collaboration: Partner with SRE, DevOps, and Platform teams to provide clear infrastructure mitigation paths. Assist in writing, reviewing, or modifying cloud configuration templates directly when necessary to resolve security flaws. - Incident Response Support: Assist Incident Response teams during active cloud or host-level breaches. Help develop and implement immediate, real-time cloud, network, or IAM configuration countermeasures to contain threats. Required Skills & Experience - Experience: 5 years of experience in Cloud Security, DevSecOps, or Systems Engineering roles. - Cloud Infrastructure Depth: Strong foundational experience working with multi-cloud environments (Deep GCP expertise preferred, with working knowledge of AWS or Azure). - Posture Management & Scanning Tooling: Hands-on experience operating modern infrastructure security platforms such as Wiz, Orca, Prisma Cloud, Lacework, or cloud-native options (GCP Security Command Center). - IaC and Automation Fluency: Strong proficiency with Infrastructure as Code platforms (Terraform, Pulumi) and GitOps deployment workflows. Ability to evaluate and configure IaC scanners like Checkov, Tfsec, or KICS. - Containerization & Orchestration: Deep understanding of Docker/container security and Kubernetes architectures (e.g., GKE, EKS), including runtime security, network policies, and workload identity. - Compliance Awareness: Understanding of how infrastructure configurations and vulnerability management map to security compliance frameworks like SOC 2, ISO 27001, CIS Benchmarks, or NIST. What We Value - Systems Thinking: The ability to see the "big picture" and understand how security decisions impact the entire stack. - Technical Influence: The ability to drive technical alignment across the organization through expertise and collaboration rather than direct authority. - Autonomy: Comfortable leading major technical initiatives and driving outcomes with minimal oversight. - Problem-Solving Mindset: A passion for breaking down complex security challenges into elegant, scalable engineering solutions. This is a full-time role that can be held from our Foster City, CA office. The role has an in-office requirement of Monday, Wednesday, and Friday. Full-Time Employee Benefits Include: πŸ’° Competitive Salary & Equity πŸ’Ή 401(k) Program with a 4% match (US Only) βš•οΈ Health, Dental, Vision and Life Insurance 🩼 Short Term and Long Term Disability 🚼 Paid Parental, Medical, Caregiver Leave 🏝 Flexible Time Off (FTO) + Holidays πŸš— Commuter Benefits (In-Office Only) πŸ“± Monthly Wellness Stipend πŸ§‘β€πŸ’» Autonomous Work Environment πŸ–₯ In Office Set-Up Reimbursement (In-Office Only) πŸš€ Quarterly Team Gatherings β˜• In Office Amenities (In-Office Only) Want to learn more about what we are up to? - Meet the Replit Agent - Replit: Make an app for that - Replit Blog - Amjad TED Talk Interviewing + Culture at Replit - Operating Principles - Reasons not to work at Replit To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.

Foster City, CA hybrid FullTime
AWSAzureDocker +3 more

Director, Forward Deployed Engineering

3mo ago
Scale AI

Scale AI

Scale AI is seeking a Director of Forward Deployed Engineering to lead delivery for their largest and most strategic enterprise accounts. This role involves embedding elite engineers with Fortune 500 customers to ship deeply integrated AI agents, overseeing everything from agent design and evaluation to the backend services, data systems, and cloud infrastructure required for production readiness. The Director will partner closely with Product and core Engineering teams to ensure efficient delivery, translating field learnings into reusable platform capabilities. This is a high-autonomy, founder-mindset position focused on owning outcomes end-to-end, scaling a high-performing team, and shaping enterprise AI adoption.

London, UK remote
AWSAzureDocker +8 more

Senior AI Infrastructure Engineer - Training Platform

3mo ago
Scale AI

Scale AI

As a Software Engineer on the Machine Learning Infrastructure team, you will build the "Operating System" for our large-scale GPU clusters. You will architect a high-performance training platform that handles the immense complexity of multi-thousand GPU workloads, ensuring every cycle is used efficiently. Your work directly determines the velocity at which our researchers can train and iterate on the world’s most advanced models. The ideal candidate is a systems expert who thrives on solving the orchestration, networking, and reliability challenges that emerge at massive scale. You will partner closely with researchers to build a seamless, resilient environment that transforms raw compute into breakthrough AI.

$216k - $270k

San Francisco, CA; Seattle, WA; New York, NY onsite
AWSKubernetesPython +8 more

Software Engineer, Robotics

3mo ago
Scale AI

Scale AI

Scale's Robotics business unit is focused on solving the data bottleneck in Physical AI across Robotics, Autonomous Vehicles, and Computer Vision. In this role, you will be a key contributor building production systems for robotics data collection, model training pipelines, and evaluation infrastructure. You will have the opportunity to own critical parts of our robotics platform, work directly with cutting-edge robotics and AV customers, and shape the future of embodied AI systems.

Mexico City, MX onsite
AWSDockerKubernetes +9 more

Staff Software Engineer, Public Sector

3mo ago
Scale AI

Scale AI

The Public Sector software engineers at Scale create core product building blocks for forward-deployed teams developing agentic capabilities across multiple domains. This role involves building systems to ingest and process federal datasets for real-time decision-making in contested environments, and developing novel agentic capabilities such as multi-layered guardrails, optimized data retrieval, and asynchronous agent orchestration. As a Staff Software Engineer, you will orchestrate the implementation of vertical features and horizontal capabilities, mentor other engineers, and define technical strategy for key agentic components. You will ensure system reliability and performance across various security classifications and network types, and communicate technical trade-offs to senior stakeholders to influence long-term product strategy.

$189k - $362k

San Francisco, CA; St. Louis, MO; New York, NY; Washington, DC remote
AWSAzureDocker +7 more

ML Systems Engineer, Robotics

3mo ago
Scale AI

Scale AI

Scale's Physical AI business unit is focused on solving data bottlenecks in Robotics, Autonomous Vehicles, and Computer Vision. This role involves applied research and developing ML pipelines for processing, training, and fine-tuning data collected by Scale, with an emphasis on optimizing algorithms and pipelines for efficient GPU execution in the cloud. You will advance research, shape Scale's offerings, and expand the frontier of data and model evaluation for Physical AI. As an ML Systems Engineer, you will design and build platforms for scalable, reliable, and efficient serving of foundation models tailored for physical agents, powering both internal research and external customer use cases.

$249k - $311k

San Francisco, CA remote
AWSDockerKubernetes +7 more

Senior Software Engineer, Public Sector

3mo ago
Scale AI

Scale AI

Scale is seeking Senior Software Engineers to join our Public Sector team. You will build core product components that enable forward-deployed teams to develop agentic capabilities across multiple domains. This involves creating systems to ingest and process federal datasets for real-time decision-making in challenging environments. You will lead the development of new agentic capabilities, including multi-layered guardrails, optimized data retrieval, orchestration of asynchronous agents, automated deviation alerts, and decision-path illustration.

$162k - $311k

San Francisco, CA; St. Louis, MO; New York, NY; Washington, DC onsite
AWSAzureDocker +11 more

Principal AI Ops Architect, GPS

3mo ago
Scale AI

Scale AI

Scale's Global Public Sector team is dedicated to leveraging AI to tackle significant challenges within the public sector worldwide. This involves creating custom AI applications impacting millions, generating high-quality training data for national LLMs, and providing AI upskilling and advisory services. As a Principal AI Ops Architect, you will be instrumental in designing and developing the production lifecycle for full-stack AI applications. Your role will encompass ensuring end-to-end system reliability, real-time inference observability, sovereign data orchestration, secure software integration, and the resilient cloud infrastructure necessary for international government partners. At Scale, we empower the public sector to transform operations and enhance citizen services through advanced technology, and we are looking for individuals ready to shape the future of AI in this domain.

Doha, Qatar; London, UK onsite
KubernetesVector DatabasesMLOps +2 more

AI Applications Ops Lead, GPS

3mo ago
Scale AI

Scale AI

Scale's Global Public Sector team is focused on using AI to address critical challenges facing the public sector worldwide. This role involves designing and developing the production lifecycle of full-stack AI applications, ensuring end-to-end system reliability, real-time inference observability, sovereign data orchestration, high-security software integration, and resilient cloud infrastructure for international government partners. The goal is to enable the public sector to transform operations and better serve citizens through cutting-edge technology, with the opportunity to be a founding member of the team.

Doha, Qatar; London, UK remote
KubernetesVector DatabasesMLOps +2 more

Staff Software Engineer, Data Platform

3mo ago
Scale AI

Scale AI

Scale is at the forefront of the AI revolution, developing data engines and technologies that power the world's leading LLMs and generative models. This role is on the Platform Engineering team, responsible for the foundational data infrastructure that supports these cutting-edge AI products. You will lead the design and development of core data storage, streaming, caching, and indexing platforms, gaining exposure to the rapidly evolving AI landscape across various industries. The work involves driving architecture, implementation, and reliability of these critical systems, collaborating with stakeholders, and mentoring junior engineers.

$252k - $315k

San Francisco, CA; New York, NY remote
KubernetesPythonFine-Tuning +10 more

Staff Infrastructure Software Engineer, Enterprise AI

3mo ago
Scale AI

Scale AI

Scale GP is seeking a Senior or Staff Infrastructure Engineer to lead the engineering of the 'paved road' for knowledge retrieval and inference engines, defining deployment standards for Agentic workflows at scale. This role bridges complex AI orchestration with world-class infrastructure, ensuring platform reliability for enterprise agents. The ideal candidate is passionate about deep technical work, mentoring, and setting long-term technical strategy while maintaining a hands-on delivery focus. You will architect and implement solutions across multiple cloud providers (GCP, Azure, AWS) for customers in regulated industries.

$252k - $315k

New York, NY; San Francisco, CA remote
AWSAzureKubernetes +13 more

Software Engineer, Enterprise AI

3mo ago
Scale AI

Scale AI

Scale Generative AI Platform (Scale GP) is an enterprise-grade platform offering APIs for knowledge retrieval, inference, and evaluation. We are seeking a skilled engineer to contribute to building and scaling our product in a dynamic environment. The role involves owning significant product areas, collaborating across backend and frontend development, and integrating with LLMs and ML models. You will tackle complex engineering challenges related to scalability and reliability, working throughout the product lifecycle from concept to production.

$216k - $270k

New York, NY; San Francisco, CA onsite
AWSAzureKubernetes +7 more

Software Engineer, Robotics & Autonomous Systems

3mo ago
Scale AI

Scale AI

Scale's Robotics business unit is dedicated to solving the data bottleneck in Physical AI across Robotics, Autonomous Vehicles, and Computer Vision. In this role, you'll be a key contributor building production systems for robotics data collection, model training pipelines, and evaluation infrastructure. You'll have the opportunity to own critical parts of our robotics platform, work directly with cutting-edge robotics and AV customers, and shape the future of embodied AI systems.

$180k - $225k

San Francisco, CA remote
AWSDockerKubernetes +9 more

Senior AI Infrastructure Engineer, Model Serving Platform

3mo ago
Scale AI

Scale AI

As a Software Engineer on the ML Infrastructure team, you will design and build platforms for scalable, reliable, and efficient serving of LLMs. Our platform powers cutting-edge research and production systems, supporting both internal and external use cases across various environments. You'll work in a highly collaborative environment, bridging research and engineering to deliver seamless experiences to our customers and accelerate innovation across the company.

$216k - $270k

San Francisco, CA; New York, NY remote
AWSDockerKubernetes +7 more

Principal Architect

3mo ago
Scale AI

Scale AI

We are seeking a Principal Architect to drive the design, development, and deployment of our agentic AI products in a fast-paced, collaborative, extremely high-visibility environment. In this role, you will lead a team of 50+ engineers, providing both strategic and technical guidance. You’ll be responsible for high-impact architectural decisions, cross-company collaboration, and executive level engagements.

$298k - $373k

Washington, DC onsite
AWSAzureKubernetes +6 more

Senior Software Engineer, Full-Stack – Scale GP

3mo ago
Scale AI

Scale AI

Scale GP (Scale Generative AI Platform) is seeking a Senior Full-Stack Engineer to help build, scale, and refine its enterprise-grade Generative AI platform. This role involves working across the full stack, from React/TypeScript frontends to Python-based backends, and integrating with LLMs and machine learning systems. The engineer will tackle complex challenges in scalability, reliability, and product experience, owning significant product areas in a fast-paced environment.

$216k - $270k

San Francisco, CA; New York, NY onsite
AWSAzureKubernetes +8 more

Infrastructure Software Engineer, Enterprise GenAI

3mo ago
Scale AI

Scale AI

Scale GP (Scale Generative AI Platform) is an enterprise-grade AI platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale our core infrastructure in a fast-paced environment. The ideal candidate will have a strong understanding of software engineering principles and practices, as well as experience with large-scale distributed systems. You will implement solutions across multiple cloud providers (GCP, Azure, AWS) for customers in diverse, highly-regulated industries like healthcare, telecom, finance, and retail.

$216k - $270k

San Francisco, CA; New York, NY remote
AWSAzureKubernetes +7 more

Software Engineer, Frontier AI Infrastructure

3mo ago
Scale AI

Scale AI

Scale AI is seeking a Software Engineer to join our dynamic Public Sector Engineering team. In this role, you will own the model inference layer, enabling state-of-the-art models, debugging AI tools, managing networking, and tracking AI model metrics. You will lead technical discussions with cloud vendors and customers, debug platform issues, and work with Product to understand features before they break. The position involves designing and implementing secure, scalable backend systems for Public Sector customers, re-architecting the stack for compliant environments, and building integration tests to prevent failures. You will also participate in customer engagements and contribute to the platform roadmap and product strategy.

$138k - $259k

San Francisco, CA; St. Louis, MO; New York, NY; Washington, DC hybrid
AWSAzureDocker +7 more

Deployed Engineer (Bay Area)

3mo ago
L

Langchain

This role focuses on working directly with companies building and running AI agents in production, helping to transform ideas and prototypes into reliable systems. It's a hands-on, technical position that partners with customer engineers throughout the entire lifecycle, from pre-sales evaluations to post-deployment advisory. The core objective is to achieve technical success, co-design agent architectures, and assist customers in operating agents reliably at scale using the LangChain suite. This position sits at the intersection of engineering, product, and go-to-market, influencing LangChain's adoption and providing valuable field insights back into the platform.

$165k - $315k

San Francisco, CA onsite FullTime
LangGraphLangChainAWS +5 more

Frontier Agent Engineering Manager, Enterprise

3mo ago
Scale AI

Scale AI

As a Forward Deployed AI Engineering Manager on our Enterprise team, you will serve as the technical bridge between Scale AI's cutting-edge AI capabilities and our most strategic customers. You will work with enterprise clients to understand their unique challenges, lead a team that architects specific AI solutions, and ensure successful deployment and adoption of AI systems in production environments. This management role combines deep engineering and AI expertise with leadership and customer-facing problem-solving, integrating AI into critical customer workflows.

San Francisco, CA; New York, NY remote
OpenAILangChainLlamaIndex +9 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

Β© 2026 AI Job Board. All rights reserved.