Kubernetes Jobs
204 open roles mentioning Kubernetes
Security Engineer - Vuln Management (Infra)
Replit
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are seeking a mid-level Infrastructure Vulnerability Management Engineer with a strong background in Cloud Security, DevSecOps, and Infrastructure-as-Code (IaC). In this role, you will bridge the gap between security, compliance, DevOps, and Platform engineering teams. You will identify infrastructure misconfigurations, secure multi-cloud environments, and manage continuous vulnerability lifecycles across cloud workloads, containers, and data repositories to satisfy strict regulatory compliance frameworks. You will also serve as a technical infrastructure responder during security incidents, deploying real-time cloud or network countermeasures to protect our production ecosystem. What You'll Do Core Responsibilities - Infrastructure Scanning & Triage: Perform continuous security scanning across our cloud posture and workloads. Review, validate, and prioritize flaws and misconfigurations based on CVSS scores, real-world exploitability, and infrastructure network exposure. - Posture Management & Visibility: Own and optimize Cloud Security Posture Management (CSPM), Kubernetes Security Posture Management (KSPM), and Data Security Posture Management (DSPM) tools to ensure uniform compliance, prevent data leakage, and maintain hardened baselines. - Infrastructure-as-Code (IaC) Security: Configure, tune, and embed automated IaC security scanning tools into CI/CD pipelines to identify architectural risks (e.g., overly permissive IAM, public S3 buckets/Cloud Storage) before they are deployed to production. - Workload & Container Security: Manage the continuous vulnerability scanning lifecycle for container images, registries, and Virtual Machines (VMs), partnering with SRE and Platform teams to build automated base-image patching and rolling upgrade pipelines. - Compliance-Driven Tracking: Track, document, and manage infrastructure vulnerabilities according to strict compliance SLAs (e.g., SOC 2, ISO 27001, PCI-DSS). Maintain audit-ready evidence of infrastructure remediation timelines and exception approvals. - Executive Reporting & Alerting: Escalate and report critical production exposures directly to the CISO and senior leadership. Maintain dashboards and alerting mechanisms that visualize infrastructure risk trends and cloud compliance posture. - Remediation Collaboration: Partner with SRE, DevOps, and Platform teams to provide clear infrastructure mitigation paths. Assist in writing, reviewing, or modifying cloud configuration templates directly when necessary to resolve security flaws. - Incident Response Support: Assist Incident Response teams during active cloud or host-level breaches. Help develop and implement immediate, real-time cloud, network, or IAM configuration countermeasures to contain threats. Required Skills & Experience - Experience: 5 years of experience in Cloud Security, DevSecOps, or Systems Engineering roles. - Cloud Infrastructure Depth: Strong foundational experience working with multi-cloud environments (Deep GCP expertise preferred, with working knowledge of AWS or Azure). - Posture Management & Scanning Tooling: Hands-on experience operating modern infrastructure security platforms such as Wiz, Orca, Prisma Cloud, Lacework, or cloud-native options (GCP Security Command Center). - IaC and Automation Fluency: Strong proficiency with Infrastructure as Code platforms (Terraform, Pulumi) and GitOps deployment workflows. Ability to evaluate and configure IaC scanners like Checkov, Tfsec, or KICS. - Containerization & Orchestration: Deep understanding of Docker/container security and Kubernetes architectures (e.g., GKE, EKS), including runtime security, network policies, and workload identity. - Compliance Awareness: Understanding of how infrastructure configurations and vulnerability management map to security compliance frameworks like SOC 2, ISO 27001, CIS Benchmarks, or NIST. What We Value - Systems Thinking: The ability to see the "big picture" and understand how security decisions impact the entire stack. - Technical Influence: The ability to drive technical alignment across the organization through expertise and collaboration rather than direct authority. - Autonomy: Comfortable leading major technical initiatives and driving outcomes with minimal oversight. - Problem-Solving Mindset: A passion for breaking down complex security challenges into elegant, scalable engineering solutions. This is a full-time role that can be held from our Foster City, CA office. The role has an in-office requirement of Monday, Wednesday, and Friday. Full-Time Employee Benefits Include: π° Competitive Salary & Equity πΉ 401(k) Program with a 4% match (US Only) βοΈ Health, Dental, Vision and Life Insurance π©Ό Short Term and Long Term Disability πΌ Paid Parental, Medical, Caregiver Leave π Flexible Time Off (FTO) + Holidays π Commuter Benefits (In-Office Only) π± Monthly Wellness Stipend π§βπ» Autonomous Work Environment π₯ In Office Set-Up Reimbursement (In-Office Only) π Quarterly Team Gatherings β In Office Amenities (In-Office Only) Want to learn more about what we are up to? - Meet the Replit Agent - Replit: Make an app for that - Replit Blog - Amjad TED Talk Interviewing + Culture at Replit - Operating Principles - Reasons not to work at Replit To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.
Staff Software Engineer, Fraud
Replit
The Fraud team is the front line defending Replit's platform from exploitation, detecting and shutting down phishing deployments, preventing cryptomining on free-tier infrastructure, stopping LLM token farming, and keeping bad actors from weaponizing the platform against our users. This is adversarial work where attackers adapt constantly, and you will build the detection systems, heuristics, and automated responses to stay ahead of them. This role offers unique experience applying AI to security problems in production, building guardrails for AI-generated code, detecting prompt injection attacks at scale, and using LLMs as a defensive tool against abuse. You will own problems end-to-end, from identifying emerging abuse patterns to shipping the systems that stop them at scale.
Staff Software Engineer, Risk
Replit
The Risk team at Replit is on the front lines, defending the platform from exploitation by detecting and shutting down malicious activities. This adversarial role involves staying ahead of constantly adapting attackers by building detection systems, heuristics, and automated responses. You will tackle unique, AI-native problems such as building guardrails for AI-generated code, detecting prompt injection attacks at scale, and leveraging LLMs as a defensive tool against abuse. This is a hands-on opportunity to apply AI to security problems in a production environment with real attackers, owning problems end-to-end from identifying abuse patterns to shipping scalable solutions.
Staff Software Engineer, Trust & Safety
Replit
Replit is seeking a Staff Software Engineer for its Trust & Safety team to defend the platform from exploitation. This role involves detecting and shutting down various forms of abuse, including phishing, cryptomining, and LLM token farming, by building advanced detection systems and automated responses. The position offers unique challenges at the intersection of AI and security, focusing on problems like building guardrails for AI-generated code, detecting prompt injection at scale, and using LLMs defensively. You will own problems end-to-end, from identifying abuse patterns to implementing solutions that stop them at scale, gaining hands-on experience applying AI to security in a production environment with real attackers.
Senior Software Engineer, Fraud
Replit
The Fraud team at Replit is on the front lines, defending the platform from exploitation by detecting and shutting down malicious activities. This adversarial role involves building detection systems, heuristics, and automated responses to combat threats like phishing, cryptomining, LLM token farming, and platform weaponization. You will work on unique AI-native security challenges, including building guardrails for AI-generated code, detecting prompt injection attacks at scale, and leveraging LLMs for defense. This is an opportunity to gain hands-on experience applying AI to security problems in a production environment with real attackers, owning problems end-to-end from identifying abuse patterns to shipping scalable solutions.
Senior Software Engineer, Risk
Replit
The Risk team at Replit is responsible for defending the platform from exploitation by detecting and shutting down malicious activities such as phishing, cryptomining, and LLM token farming. This role involves adversarial work, where you will build detection systems, heuristics, and automated responses to stay ahead of constantly adapting attackers. A unique aspect of this role is the AI-native nature of Replit's platform, offering hands-on experience applying AI to security problems in a production environment. You will own problems end-to-end, from identifying emerging abuse patterns to shipping systems that stop them at scale, working on novel problems like building guardrails for AI-generated code and detecting prompt injection attacks.
Senior Software Engineer, Trust & Safety
Replit
The Trust & Safety team at Replit is on the front lines, defending the platform from exploitation by detecting and shutting down malicious activities such as phishing, cryptomining, and LLM token farming. This role involves adversarial work, requiring the development of detection systems, heuristics, and automated responses to stay ahead of constantly adapting attackers. A unique aspect of this position is its focus on AI-native security challenges, including building guardrails for AI-generated code, detecting prompt injection attacks at scale, and leveraging LLMs for defense against platform abuse. The role offers hands-on experience applying AI to security problems in a production environment with real attackers, with opportunities to own problems end-to-end from identifying abuse patterns to shipping scalable solutions.
Director, Forward Deployed Engineering
Scale AI
Scale AI is seeking a Director of Forward Deployed Engineering to lead delivery for their largest and most strategic enterprise accounts. This role involves embedding elite engineers with Fortune 500 customers to ship deeply integrated AI agents, overseeing everything from agent design and evaluation to the backend services, data systems, and cloud infrastructure required for production readiness. The Director will partner closely with Product and core Engineering teams to ensure efficient delivery, translating field learnings into reusable platform capabilities. This is a high-autonomy, founder-mindset position focused on owning outcomes end-to-end, scaling a high-performing team, and shaping enterprise AI adoption.
Senior AI Infrastructure Engineer - Training Platform
Scale AI
As a Software Engineer on the Machine Learning Infrastructure team, you will build the "Operating System" for our large-scale GPU clusters. You will architect a high-performance training platform that handles the immense complexity of multi-thousand GPU workloads, ensuring every cycle is used efficiently. Your work directly determines the velocity at which our researchers can train and iterate on the worldβs most advanced models. The ideal candidate is a systems expert who thrives on solving the orchestration, networking, and reliability challenges that emerge at massive scale. You will partner closely with researchers to build a seamless, resilient environment that transforms raw compute into breakthrough AI.
$216k - $270k
Senior Software Engineer, Public Sector
Scale AI
Scale is seeking Senior Software Engineers to join our Public Sector team. You will build core product components that enable forward-deployed teams to develop agentic capabilities across multiple domains. This involves creating systems to ingest and process federal datasets for real-time decision-making in challenging environments. You will lead the development of new agentic capabilities, including multi-layered guardrails, optimized data retrieval, orchestration of asynchronous agents, automated deviation alerts, and decision-path illustration.
$162k - $311k
Software Engineer, Robotics
Scale AI
Scale's Robotics business unit is focused on solving the data bottleneck in Physical AI across Robotics, Autonomous Vehicles, and Computer Vision. In this role, you will be a key contributor building production systems for robotics data collection, model training pipelines, and evaluation infrastructure. You will have the opportunity to own critical parts of our robotics platform, work directly with cutting-edge robotics and AV customers, and shape the future of embodied AI systems.
Staff Software Engineer, Public Sector
Scale AI
The Public Sector software engineers at Scale create core product building blocks for forward-deployed teams developing agentic capabilities across multiple domains. This role involves building systems to ingest and process federal datasets for real-time decision-making in contested environments, and developing novel agentic capabilities such as multi-layered guardrails, optimized data retrieval, and asynchronous agent orchestration. As a Staff Software Engineer, you will orchestrate the implementation of vertical features and horizontal capabilities, mentor other engineers, and define technical strategy for key agentic components. You will ensure system reliability and performance across various security classifications and network types, and communicate technical trade-offs to senior stakeholders to influence long-term product strategy.
$189k - $362k
ML Systems Engineer, Robotics
Scale AI
Scale's Physical AI business unit is focused on solving data bottlenecks in Robotics, Autonomous Vehicles, and Computer Vision. This role involves applied research and developing ML pipelines for processing, training, and fine-tuning data collected by Scale, with an emphasis on optimizing algorithms and pipelines for efficient GPU execution in the cloud. You will advance research, shape Scale's offerings, and expand the frontier of data and model evaluation for Physical AI. As an ML Systems Engineer, you will design and build platforms for scalable, reliable, and efficient serving of foundation models tailored for physical agents, powering both internal research and external customer use cases.
$249k - $311k
Principal AI Ops Architect, GPS
Scale AI
Scale's Global Public Sector team is dedicated to leveraging AI to tackle significant challenges within the public sector worldwide. This involves creating custom AI applications impacting millions, generating high-quality training data for national LLMs, and providing AI upskilling and advisory services. As a Principal AI Ops Architect, you will be instrumental in designing and developing the production lifecycle for full-stack AI applications. Your role will encompass ensuring end-to-end system reliability, real-time inference observability, sovereign data orchestration, secure software integration, and the resilient cloud infrastructure necessary for international government partners. At Scale, we empower the public sector to transform operations and enhance citizen services through advanced technology, and we are looking for individuals ready to shape the future of AI in this domain.
Staff Software Engineer, Data Platform
Scale AI
Scale is at the forefront of the AI revolution, developing data engines and technologies that power the world's leading LLMs and generative models. This role is on the Platform Engineering team, responsible for the foundational data infrastructure that supports these cutting-edge AI products. You will lead the design and development of core data storage, streaming, caching, and indexing platforms, gaining exposure to the rapidly evolving AI landscape across various industries. The work involves driving architecture, implementation, and reliability of these critical systems, collaborating with stakeholders, and mentoring junior engineers.
$252k - $315k
AI Applications Ops Lead, GPS
Scale AI
Scale's Global Public Sector team is focused on using AI to address critical challenges facing the public sector worldwide. This role involves designing and developing the production lifecycle of full-stack AI applications, ensuring end-to-end system reliability, real-time inference observability, sovereign data orchestration, high-security software integration, and resilient cloud infrastructure for international government partners. The goal is to enable the public sector to transform operations and better serve citizens through cutting-edge technology, with the opportunity to be a founding member of the team.
Senior AI Infrastructure Engineer, Model Serving Platform
Scale AI
As a Software Engineer on the ML Infrastructure team, you will design and build platforms for scalable, reliable, and efficient serving of LLMs. Our platform powers cutting-edge research and production systems, supporting both internal and external use cases across various environments. You'll work in a highly collaborative environment, bridging research and engineering to deliver seamless experiences to our customers and accelerate innovation across the company.
$216k - $270k
Software Engineer, Enterprise AI
Scale AI
Scale Generative AI Platform (Scale GP) is an enterprise-grade platform offering APIs for knowledge retrieval, inference, and evaluation. We are seeking a skilled engineer to contribute to building and scaling our product in a dynamic environment. The role involves owning significant product areas, collaborating across backend and frontend development, and integrating with LLMs and ML models. You will tackle complex engineering challenges related to scalability and reliability, working throughout the product lifecycle from concept to production.
$216k - $270k
Senior Software Engineer, Full-Stack β Scale GP
Scale AI
Scale GP (Scale Generative AI Platform) is seeking a Senior Full-Stack Engineer to help build, scale, and refine its enterprise-grade Generative AI platform. This role involves working across the full stack, from React/TypeScript frontends to Python-based backends, and integrating with LLMs and machine learning systems. The engineer will tackle complex challenges in scalability, reliability, and product experience, owning significant product areas in a fast-paced environment.
$216k - $270k
Principal Architect
Scale AI
We are seeking a Principal Architect to drive the design, development, and deployment of our agentic AI products in a fast-paced, collaborative, extremely high-visibility environment. In this role, you will lead a team of 50+ engineers, providing both strategic and technical guidance. Youβll be responsible for high-impact architectural decisions, cross-company collaboration, and executive level engagements.
$298k - $373k