Terraform Jobs
187 open roles mentioning Terraform
DevOps Engineer - AI Systems
Scale AI
We are seeking a skilled DevOps Engineer to manage our cloud infrastructure, specifically for AI training and inference workloads. This role will involve working with major cloud providers and utilizing containerization and infrastructure-as-code tools to ensure efficient and scalable operations.
$145k - $195k
Technical Program Manager, Platform & Infrastructure
Harvey
Harvey is transforming legal and professional services by integrating frontier agentic AI with an enterprise-grade platform. We are seeking a Staff Technical Program Manager, Platform & Infrastructure to lead critical scaling initiatives. This role involves acting as the central point of contact across Core Infrastructure, Backend Platform, and product engineering teams, as well as cross-functional departments like Security and Finance. You will be responsible for multi-quarter horizontal programs focused on cloud migrations, cost optimization, capacity planning, BYOC, and building foundational internal platform elements. The position requires deep technical engagement with senior engineers, aligning leadership on long-term plans, and driving programs from initial concept to successful delivery.
$188k - $278k
Staff Infrastructure Engineer
Replit
Replit is seeking a Staff Infrastructure Engineer to join their Infrastructure Engineering team. This role focuses on ensuring the reliability, scalability, and performance of Replit's platform, which serves millions of developers globally. The engineer will bridge development and operations, implementing automation and best practices to enable efficient scaling and high availability. The position involves proactively identifying and resolving reliability issues, designing robust monitoring solutions, automating operational tasks, and mentoring the engineering team on reliability principles.
Applied AI Engineer, Site Reliability Engineer - EMEA
Mistral AI
Mistral AI is seeking a founding engineer for its Applied AI Site Reliability Engineering (SRE) sub-team. This role is crucial for building and operating a framework that ensures the reliability and sustainability of Mistral's AI solutions across all customer accounts, whether hosted by Mistral or the customer. You will operate in four key modes: BUILD (designing for a fleet of platforms, proactive reliability, authoring runbooks, implementing observability), RUN (operating Tier-1 customer environments, ensuring SLO compliance, managing incidents), ENABLE (productizing deployment, security, and scaling of Applied AI solutions), and SECURE (owning security operations, leading CVE response, and implementing supply-chain integrity controls). This is a framework-first, fleet management role focused on structurally solving problems for all customers, not just individual ones. The team values people and outputs, direct feedback, low ego, and high standards in a fast-paced, unstructured environment.
Software Engineer (Backend), Enterprise
Scale AI
Scale AI is seeking a Backend Engineer to join their team and build the core infrastructure for large-scale GenAI systems. This role involves designing and implementing scalable APIs, distributed data systems, and robust deployment pipelines to ensure production-grade reliability and performance for enterprise AI products. You will be instrumental in shaping how AI systems are deployed and scaled in the real world, working at the forefront of the GenAI revolution and solving complex backend and infrastructure challenges. This is an opportunity to contribute to cutting-edge solutions that transform workflows and drive efficiency for major enterprises.
Software Engineer, Enterprise
Scale AI
Scale AI is pioneering the next era of enterprise AI, providing cutting-edge solutions that transform workflows and automate complex processes for large enterprises. The Scale Generative AI Platform (SGP) offers foundational services and APIs for seamless AI integration at production scale. This role focuses on building the core infrastructure for large-scale GenAI systems, designing scalable APIs, distributed data systems, and robust deployment pipelines to ensure production-grade reliability and performance. It's an opportunity to solve hard backend and infrastructure challenges that enable AI to work at enterprise scale and shape how AI systems are deployed and scaled in the real world.
Senior Cloud Infrastructure Engineer
Langfuse
Langfuse is seeking a Senior Cloud Infrastructure Engineer to ensure the reliability, performance, and cost-efficiency of their open-source LLM engineering platform. This role is crucial for maintaining Langfuse Cloud on AWS ECS Fargate and ClickHouse Cloud, as well as supporting self-hosted deployments. The engineer will own the end-to-end observability stack, automate infrastructure processes, and scale the platform to meet growing demand. This is an opportunity to work closely with the ClickHouse team and directly impact a product used by major enterprises.
Security Engineer - Vuln Management (Infra)
Replit
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are seeking a mid-level Infrastructure Vulnerability Management Engineer with a strong background in Cloud Security, DevSecOps, and Infrastructure-as-Code (IaC). In this role, you will bridge the gap between security, compliance, DevOps, and Platform engineering teams. You will identify infrastructure misconfigurations, secure multi-cloud environments, and manage continuous vulnerability lifecycles across cloud workloads, containers, and data repositories to satisfy strict regulatory compliance frameworks. You will also serve as a technical infrastructure responder during security incidents, deploying real-time cloud or network countermeasures to protect our production ecosystem. What You'll Do Core Responsibilities - Infrastructure Scanning & Triage: Perform continuous security scanning across our cloud posture and workloads. Review, validate, and prioritize flaws and misconfigurations based on CVSS scores, real-world exploitability, and infrastructure network exposure. - Posture Management & Visibility: Own and optimize Cloud Security Posture Management (CSPM), Kubernetes Security Posture Management (KSPM), and Data Security Posture Management (DSPM) tools to ensure uniform compliance, prevent data leakage, and maintain hardened baselines. - Infrastructure-as-Code (IaC) Security: Configure, tune, and embed automated IaC security scanning tools into CI/CD pipelines to identify architectural risks (e.g., overly permissive IAM, public S3 buckets/Cloud Storage) before they are deployed to production. - Workload & Container Security: Manage the continuous vulnerability scanning lifecycle for container images, registries, and Virtual Machines (VMs), partnering with SRE and Platform teams to build automated base-image patching and rolling upgrade pipelines. - Compliance-Driven Tracking: Track, document, and manage infrastructure vulnerabilities according to strict compliance SLAs (e.g., SOC 2, ISO 27001, PCI-DSS). Maintain audit-ready evidence of infrastructure remediation timelines and exception approvals. - Executive Reporting & Alerting: Escalate and report critical production exposures directly to the CISO and senior leadership. Maintain dashboards and alerting mechanisms that visualize infrastructure risk trends and cloud compliance posture. - Remediation Collaboration: Partner with SRE, DevOps, and Platform teams to provide clear infrastructure mitigation paths. Assist in writing, reviewing, or modifying cloud configuration templates directly when necessary to resolve security flaws. - Incident Response Support: Assist Incident Response teams during active cloud or host-level breaches. Help develop and implement immediate, real-time cloud, network, or IAM configuration countermeasures to contain threats. Required Skills & Experience - Experience: 5 years of experience in Cloud Security, DevSecOps, or Systems Engineering roles. - Cloud Infrastructure Depth: Strong foundational experience working with multi-cloud environments (Deep GCP expertise preferred, with working knowledge of AWS or Azure). - Posture Management & Scanning Tooling: Hands-on experience operating modern infrastructure security platforms such as Wiz, Orca, Prisma Cloud, Lacework, or cloud-native options (GCP Security Command Center). - IaC and Automation Fluency: Strong proficiency with Infrastructure as Code platforms (Terraform, Pulumi) and GitOps deployment workflows. Ability to evaluate and configure IaC scanners like Checkov, Tfsec, or KICS. - Containerization & Orchestration: Deep understanding of Docker/container security and Kubernetes architectures (e.g., GKE, EKS), including runtime security, network policies, and workload identity. - Compliance Awareness: Understanding of how infrastructure configurations and vulnerability management map to security compliance frameworks like SOC 2, ISO 27001, CIS Benchmarks, or NIST. What We Value - Systems Thinking: The ability to see the "big picture" and understand how security decisions impact the entire stack. - Technical Influence: The ability to drive technical alignment across the organization through expertise and collaboration rather than direct authority. - Autonomy: Comfortable leading major technical initiatives and driving outcomes with minimal oversight. - Problem-Solving Mindset: A passion for breaking down complex security challenges into elegant, scalable engineering solutions. This is a full-time role that can be held from our Foster City, CA office. The role has an in-office requirement of Monday, Wednesday, and Friday. Full-Time Employee Benefits Include: π° Competitive Salary & Equity πΉ 401(k) Program with a 4% match (US Only) βοΈ Health, Dental, Vision and Life Insurance π©Ό Short Term and Long Term Disability πΌ Paid Parental, Medical, Caregiver Leave π Flexible Time Off (FTO) + Holidays π Commuter Benefits (In-Office Only) π± Monthly Wellness Stipend π§βπ» Autonomous Work Environment π₯ In Office Set-Up Reimbursement (In-Office Only) π Quarterly Team Gatherings β In Office Amenities (In-Office Only) Want to learn more about what we are up to? - Meet the Replit Agent - Replit: Make an app for that - Replit Blog - Amjad TED Talk Interviewing + Culture at Replit - Operating Principles - Reasons not to work at Replit To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.
Senior AI Infrastructure Engineer - Training Platform
Scale AI
As a Software Engineer on the Machine Learning Infrastructure team, you will build the "Operating System" for our large-scale GPU clusters. You will architect a high-performance training platform that handles the immense complexity of multi-thousand GPU workloads, ensuring every cycle is used efficiently. Your work directly determines the velocity at which our researchers can train and iterate on the worldβs most advanced models. The ideal candidate is a systems expert who thrives on solving the orchestration, networking, and reliability challenges that emerge at massive scale. You will partner closely with researchers to build a seamless, resilient environment that transforms raw compute into breakthrough AI.
$216k - $270k
ML Systems Engineer, Robotics
Scale AI
Scale's Physical AI business unit is focused on solving data bottlenecks in Robotics, Autonomous Vehicles, and Computer Vision. This role involves applied research and developing ML pipelines for processing, training, and fine-tuning data collected by Scale, with an emphasis on optimizing algorithms and pipelines for efficient GPU execution in the cloud. You will advance research, shape Scale's offerings, and expand the frontier of data and model evaluation for Physical AI. As an ML Systems Engineer, you will design and build platforms for scalable, reliable, and efficient serving of foundation models tailored for physical agents, powering both internal research and external customer use cases.
$249k - $311k
Senior AI Infrastructure Engineer, Model Serving Platform
Scale AI
As a Software Engineer on the ML Infrastructure team, you will design and build platforms for scalable, reliable, and efficient serving of LLMs. Our platform powers cutting-edge research and production systems, supporting both internal and external use cases across various environments. You'll work in a highly collaborative environment, bridging research and engineering to deliver seamless experiences to our customers and accelerate innovation across the company.
$216k - $270k
Staff Infrastructure Software Engineer, Enterprise AI
Scale AI
Scale GP is seeking a Senior or Staff Infrastructure Engineer to lead the engineering of the 'paved road' for knowledge retrieval and inference engines, defining deployment standards for Agentic workflows at scale. This role bridges complex AI orchestration with world-class infrastructure, ensuring platform reliability for enterprise agents. The ideal candidate is passionate about deep technical work, mentoring, and setting long-term technical strategy while maintaining a hands-on delivery focus. You will architect and implement solutions across multiple cloud providers (GCP, Azure, AWS) for customers in regulated industries.
$252k - $315k
Software Engineer, Backend (Warsaw)
Mistral AI
Mistral AI is seeking passionate and skilled software engineers to join our Warsaw-based Context Engine team. You will be instrumental in building and enhancing Mistral's agent harness, which is crucial for delivering production-grade knowledge processing and directly influences how users interact with our AI platform. The role involves developing reliable, high-performance backend systems and APIs that serve millions of users and developers, ultimately making the user and developer experience more engaging, efficient, and intuitive. We are looking for engineers at mid to senior staff levels who are eager to learn, collaborate, and ship impactful products.
Infrastructure Engineer, Security
thinkingmachines
Thinking Machines Lab is seeking an infrastructure engineer to lead and enhance the security infrastructure for their foundation models. This role involves working across compute, storage, networking, and data platforms to ensure systems are secure, reliable, and scalable. The engineer will define security controls, architecture, and tooling, integrating security by default into the platform. Collaboration with research and product teams will be key to enabling rapid progress while maintaining robust protection for models, data, and environments.
$200k - $475k
Staff Software Engineer, Developer Experience
Harvey
Harvey is transforming legal and professional services by combining agentic AI, an enterprise-grade platform, and deep domain expertise. We are a fast-scaling company with strong product-market fit, seeking individuals who want to do the best work of their careers. The Developer Experience team is crucial for maintaining our rapid growth, building CI/CD systems, internal frameworks, and platform guardrails that enable all engineers to ship AI features quickly and safely. As a Software Engineer on this team, you will create systems to maximize engineer velocity and efficiency, working across the tech stack to instill reliability, automation, and simplicity into developer workflows. You'll also have a unique opportunity to apply AI to developer productivity, shaping the future of software development within an AI-native company.
$231k - $340k
Senior Front-End (React) Engineer
Nabla
Nabla is seeking a Senior Front-End Engineer to join a cross-functional squad focused on developing high-impact user experiences across web, desktop, browser extension, and mobile interfaces. This role involves setting UI quality and performance standards, evolving the design system, architecture, and tooling. You will lead the development of complex front-end features, drive technical decisions, partner with Design to create intuitive experiences, and collaborate with Product, ML, and Back-End teams to ship full-stack features. The position also includes mentoring other engineers and contributing to team best practices. Nabla is an AI company dedicated to improving healthcare by streamlining clinical documentation and workflows, backed by significant funding and led by experienced AI engineers.
Staff Software Engineer, Core Infrastructure
Harvey
Harvey is transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. This is a rare chance to help build a generational company at a true inflection point, scaling fast and defining a new category. The work is ambitious, the bar is high, and the opportunity for growth is unmatched. As a Staff Software Engineer on the Core Infrastructure team, you will play a critical role in designing and building new infrastructure systems while scaling and strengthening existing ones. Our infrastructure powers every user interaction with Harvey, processing billions of prompt tokens and millions of daily requests across our global legal AI platform. You'll work in an environment balanced between innovation and operational excellence, ensuring Harvey remains resilient and efficient as it scales products, regions, customers, and usage. Your contributions will directly impact the reliability, scalability, and security of our platform.
$201k - $264k
Staff Software Engineer, Core Infrastructure
Harvey
Harvey is transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. This is a rare chance to help build a generational company at a true inflection point, scaling fast and defining a new category. The work is ambitious, the bar is high, and the opportunity for growth is unmatched. The team moves fast, takes ownership, and is deeply committed to the mission, operating with intensity and pushing for excellence. As a Staff Software Engineer on the Core Infrastructure team, you will play a critical role in designing and building new infrastructure systems while scaling and strengthening existing ones. Your contributions will directly impact the reliability, scalability, and security of our platform as we serve the world's leading law firms and professional service providers.
$236k - $290k
Deployed Infrastructure Engineer
sierra.ai
Sierra is building a platform to enable companies to create better, more human customer experiences with AI. As a Forward Deployed Infrastructure Engineer, you will be a founding member responsible for the end-to-end lifecycle of customer deployments. You will work directly with enterprise clients to architect, deploy, upgrade, and manage Sierra's AI platform infrastructure, ensuring it meets their specific security, compliance, and operational needs. This role involves defining scalable and repeatable deployment standards, tooling, and processes, and collaborating closely with Engineering, Product, Sales, and Customer Success teams to manage complex customer engagements.
Data Center Controls Network Engineer
OpenAI
OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Controls Network Engineer, you will design, validate, and scale the controls and OT network architectures that support high-density AI data centers. You will work across controls systems, OT infrastructure, telemetry, commissioning, deployment, and operations, partnering with mechanical, electrical, IT/networking, security, and external delivery teams. We are seeking a mid to senior OT Network Engineer with a strong controls systems background to lead the design and operation of resilient, secure, and scalable OT network architectures for high-density AI data centers. This role translates compute, power, cooling, and operational requirements into practical OT network designs, evaluates vendor solutions, and drives technical decisions across controls infrastructure, telemetry, commissioning, and operations. The ideal candidate has strong hands-on experience in mission-critical OT environments, including industrial networking, virtualized infrastructure, and OT network operations, with expertise in routing, switching, segmentation, firewall policy, time synchronization, monitoring, and network lifecycle support.