Terraform Jobs

74 open roles mentioning Terraform

Senior Software Engineer - Together Cloud Infrastructure

1mo ago
Together AI

Together AI

Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. As a Senior AI Infrastructure Engineer, you will play a key role in building the next generation AI cloud platform – a highly available, global, blazing-fast cloud infrastructure that virtualizes cutting-edge ML hardware and enables state-of-the-art ML practitioners with self-serve AI cloud services. This platform serves both our internal SaaS products and our external cloud customers, spanning dozens of data centers across the world.

$160k - $230k

San Francisco remote
AWSAzureKubernetes +12 more

Machine Learning, Platform Engineer

1mo ago
Together AI

Together AI

Together AI is a research-driven artificial intelligence company focused on lowering the cost of modern AI systems. This role is part of a team dedicated to enabling custom models and dedicated inference on Together's platform. The team is responsible for building a container platform, optimizing autoscaling, minimizing cold starts, achieving the best end-to-end model performance, and providing a best-in-class developer experience with great tooling. The work often involves video or audio generation across the stack, including CUDA kernels, PyTorch optimization, inference engines, container orchestration, and queueing theory.

$160k - $250k

San Francisco remote
KubernetesPythonRust +7 more

AI infrastructure Engineer (SRE) Amsterdam

1mo ago
Together AI

Together AI

Together AI is seeking an AI Infrastructure Engineer (SRE) to ensure the smooth operation of user-facing services and production systems. This role combines the skills of a pragmatic operator and a software engineer, applying sound engineering principles, operational discipline, and automation to our operating environments and codebase. You will specialize in systems such as operating systems, storage subsystems, and networking, while implementing best practices for availability, reliability, and scalability, with interests in algorithms and distributed systems. Join a research-driven artificial intelligence company focused on advancing AI through open and transparent systems, aiming to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models.

Amsterdam onsite
KubernetesTerraformObservability +7 more

AI Infrastructure Engineer

1mo ago
Together AI

Together AI

As an AI Infrastructure Engineer at Together, you will be responsible for ensuring the smooth operation of all user-facing services and production systems. This role blends pragmatic operations with software engineering, applying sound engineering principles, operational discipline, and mature automation to our operating environments and codebase. You will specialize in systems (operating systems, storage subsystems, networking), implementing best practices for availability, reliability, and scalability, with varied interests in algorithms and distributed systems. Together AI is a research-driven artificial intelligence company focused on lowering the cost of modern AI systems through co-designing software, hardware, algorithms, and models, and we invite you to join our passionate group of researchers and engineers in building the next generation of AI infrastructure.

$190k - $270k

San Francisco remote
KubernetesPythonTerraform +2 more

Senior Software Engineer Together Cloud Infrastructure

1mo ago
Together AI

Together AI

Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. As a Senior AI Infrastructure Engineer, you will play a key role in building the next generation AI cloud platform – a highly available, global, blazing-fast cloud infrastructure that virtualizes cutting-edge ML hardware and enables state-of-the-art ML practitioners with self-serve AI cloud services. This platform serves both our internal SaaS products and our external cloud customers, spanning dozens of data centers across the world.

Amsterdam hybrid
AWSAzureKubernetes +12 more

DevOps Engineer, GPS

1mo ago
Scale AI

Scale AI

Scale's Global Public Sector team is focused on using AI to address critical challenges facing the public sector worldwide. We create custom AI applications impacting millions of citizens, generate high-quality training data for national LLMs, and provide upskilling and advisory services to spread AI's impact. As a DevOps Engineer, you will design and develop core platforms and software systems, supporting orchestration, data abstraction, data pipelines, identity & access management, security tools, and underlying cloud infrastructure. We are enabling the public sector to transform operations and better serve citizens through cutting-edge technology.

Dubai, UAE; Riyadh, Saudi Arabia remote
AWSAzureDocker +7 more

Software Engineer, Platform

1mo ago
Scale AI

Scale AI

Scale is at the forefront of the AI revolution, building the Generative AI Data Engine and other products that power the world's most advanced LLMs. The Platform Engineering team is foundational to these efforts, responsible for designing and developing shared platforms, architecting core cloud infrastructure, and redefining software development processes. This role offers exposure to the cutting edge of AI development across various sectors, from startups to governments. You will drive the design and implementation of critical platforms, collaborate with cross-functional teams, and proactively improve engineering practices. This is an opportunity to shape the future of AI infrastructure and contribute to some of the most important work in how humanity interacts with AI.

$216k - $270k

San Francisco, CA; New York, NY remote
AWSDockerKubernetes +11 more

DevOps Engineer - AI Systems

1mo ago
Scale AI

Scale AI

We are seeking a skilled DevOps Engineer to manage our cloud infrastructure, specifically for AI training and inference workloads. This role will involve working with major cloud providers and utilizing containerization and infrastructure-as-code tools to ensure efficient and scalable operations.

$145k - $195k

Remote remote Full-time
AWSDockerKubernetes +2 more

Technical Program Manager, Platform & Infrastructure

1mo ago
Harvey

Harvey

Harvey is transforming legal and professional services by integrating frontier agentic AI with an enterprise-grade platform. We are seeking a Staff Technical Program Manager, Platform & Infrastructure to lead critical scaling initiatives. This role involves acting as the central point of contact across Core Infrastructure, Backend Platform, and product engineering teams, as well as cross-functional departments like Security and Finance. You will be responsible for multi-quarter horizontal programs focused on cloud migrations, cost optimization, capacity planning, BYOC, and building foundational internal platform elements. The position requires deep technical engagement with senior engineers, aligning leadership on long-term plans, and driving programs from initial concept to successful delivery.

$188k - $278k

San Francisco hybrid FullTime
OpenAIAnthropicAWS +4 more

Staff Infrastructure Engineer

2mo ago
Replit

Replit

Replit is seeking a Staff Infrastructure Engineer to join their Infrastructure Engineering team. This role focuses on ensuring the reliability, scalability, and performance of Replit's platform, which serves millions of developers globally. The engineer will bridge development and operations, implementing automation and best practices to enable efficient scaling and high availability. The position involves proactively identifying and resolving reliability issues, designing robust monitoring solutions, automating operational tasks, and mentoring the engineering team on reliability principles.

Foster City, CA hybrid FullTime
DockerKubernetesPython +3 more

Applied AI Engineer, Site Reliability Engineer - EMEA

2mo ago
Mistral AI

Mistral AI

Mistral AI is seeking a founding engineer for its Applied AI Site Reliability Engineering (SRE) sub-team. This role is crucial for building and operating a framework that ensures the reliability and sustainability of Mistral's AI solutions across all customer accounts, whether hosted by Mistral or the customer. You will operate in four key modes: BUILD (designing for a fleet of platforms, proactive reliability, authoring runbooks, implementing observability), RUN (operating Tier-1 customer environments, ensuring SLO compliance, managing incidents), ENABLE (productizing deployment, security, and scaling of Applied AI solutions), and SECURE (owning security operations, leading CVE response, and implementing supply-chain integrity controls). This is a framework-first, fleet management role focused on structurally solving problems for all customers, not just individual ones. The team values people and outputs, direct feedback, low ego, and high standards in a fast-paced, unstructured environment.

Paris remote Full-time
MistralAWSAzure +12 more

Software Engineer (Backend), Enterprise

2mo ago
Scale AI

Scale AI

Scale AI is seeking a Backend Engineer to join their team and build the core infrastructure for large-scale GenAI systems. This role involves designing and implementing scalable APIs, distributed data systems, and robust deployment pipelines to ensure production-grade reliability and performance for enterprise AI products. You will be instrumental in shaping how AI systems are deployed and scaled in the real world, working at the forefront of the GenAI revolution and solving complex backend and infrastructure challenges. This is an opportunity to contribute to cutting-edge solutions that transform workflows and drive efficiency for major enterprises.

Budapest, Hungary remote
AWSAzureDocker +9 more

Software Engineer, Enterprise

2mo ago
Scale AI

Scale AI

Scale AI is pioneering the next era of enterprise AI, providing cutting-edge solutions that transform workflows and automate complex processes for large enterprises. The Scale Generative AI Platform (SGP) offers foundational services and APIs for seamless AI integration at production scale. This role focuses on building the core infrastructure for large-scale GenAI systems, designing scalable APIs, distributed data systems, and robust deployment pipelines to ensure production-grade reliability and performance. It's an opportunity to solve hard backend and infrastructure challenges that enable AI to work at enterprise scale and shape how AI systems are deployed and scaled in the real world.

London, UK remote
AWSAzureDocker +9 more

Senior Site Reliability Engineer

2mo ago
Harvey

Harvey

As a Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability, scalability, and performance of our legal AI platform. You’ll join a high-leverage team that sits at the intersection of infrastructure and product, owning the systems that keep our platform fast, secure, and always on. From scaling across 50+ regions to automating mission-critical operations, your work will ensure that Harvey remains resilient as we grow. If you’re passionate about building robust systems and reducing complexity through automation, we’d love to work with you.

Bengaluru hybrid FullTime
AWSAzureKubernetes +4 more

Infrastructure Security Engineer (Secret + Clearance)

2mo ago
Cohere

Cohere

Cohere is a leading enterprise AI company focused on building cutting-edge foundation AI models and end-to-end products. We are looking for passionate individuals to join our team and contribute to the widespread adoption of AI by training and deploying frontier models for enterprises. We are a global technology company co-headquartered in Toronto and San Francisco, with key offices in London, New York City, Montreal, Seoul, Germany, and Paris.

Toronto remote FullTime
CohereAWSAzure +3 more

Security Engineer - Vuln Management (Infra)

2mo ago
Replit

Replit

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are seeking a mid-level Infrastructure Vulnerability Management Engineer with a strong background in Cloud Security, DevSecOps, and Infrastructure-as-Code (IaC). In this role, you will bridge the gap between security, compliance, DevOps, and Platform engineering teams. You will identify infrastructure misconfigurations, secure multi-cloud environments, and manage continuous vulnerability lifecycles across cloud workloads, containers, and data repositories to satisfy strict regulatory compliance frameworks. You will also serve as a technical infrastructure responder during security incidents, deploying real-time cloud or network countermeasures to protect our production ecosystem. What You'll Do Core Responsibilities - Infrastructure Scanning & Triage: Perform continuous security scanning across our cloud posture and workloads. Review, validate, and prioritize flaws and misconfigurations based on CVSS scores, real-world exploitability, and infrastructure network exposure. - Posture Management & Visibility: Own and optimize Cloud Security Posture Management (CSPM), Kubernetes Security Posture Management (KSPM), and Data Security Posture Management (DSPM) tools to ensure uniform compliance, prevent data leakage, and maintain hardened baselines. - Infrastructure-as-Code (IaC) Security: Configure, tune, and embed automated IaC security scanning tools into CI/CD pipelines to identify architectural risks (e.g., overly permissive IAM, public S3 buckets/Cloud Storage) before they are deployed to production. - Workload & Container Security: Manage the continuous vulnerability scanning lifecycle for container images, registries, and Virtual Machines (VMs), partnering with SRE and Platform teams to build automated base-image patching and rolling upgrade pipelines. - Compliance-Driven Tracking: Track, document, and manage infrastructure vulnerabilities according to strict compliance SLAs (e.g., SOC 2, ISO 27001, PCI-DSS). Maintain audit-ready evidence of infrastructure remediation timelines and exception approvals. - Executive Reporting & Alerting: Escalate and report critical production exposures directly to the CISO and senior leadership. Maintain dashboards and alerting mechanisms that visualize infrastructure risk trends and cloud compliance posture. - Remediation Collaboration: Partner with SRE, DevOps, and Platform teams to provide clear infrastructure mitigation paths. Assist in writing, reviewing, or modifying cloud configuration templates directly when necessary to resolve security flaws. - Incident Response Support: Assist Incident Response teams during active cloud or host-level breaches. Help develop and implement immediate, real-time cloud, network, or IAM configuration countermeasures to contain threats. Required Skills & Experience - Experience: 5 years of experience in Cloud Security, DevSecOps, or Systems Engineering roles. - Cloud Infrastructure Depth: Strong foundational experience working with multi-cloud environments (Deep GCP expertise preferred, with working knowledge of AWS or Azure). - Posture Management & Scanning Tooling: Hands-on experience operating modern infrastructure security platforms such as Wiz, Orca, Prisma Cloud, Lacework, or cloud-native options (GCP Security Command Center). - IaC and Automation Fluency: Strong proficiency with Infrastructure as Code platforms (Terraform, Pulumi) and GitOps deployment workflows. Ability to evaluate and configure IaC scanners like Checkov, Tfsec, or KICS. - Containerization & Orchestration: Deep understanding of Docker/container security and Kubernetes architectures (e.g., GKE, EKS), including runtime security, network policies, and workload identity. - Compliance Awareness: Understanding of how infrastructure configurations and vulnerability management map to security compliance frameworks like SOC 2, ISO 27001, CIS Benchmarks, or NIST. What We Value - Systems Thinking: The ability to see the "big picture" and understand how security decisions impact the entire stack. - Technical Influence: The ability to drive technical alignment across the organization through expertise and collaboration rather than direct authority. - Autonomy: Comfortable leading major technical initiatives and driving outcomes with minimal oversight. - Problem-Solving Mindset: A passion for breaking down complex security challenges into elegant, scalable engineering solutions. This is a full-time role that can be held from our Foster City, CA office. The role has an in-office requirement of Monday, Wednesday, and Friday. Full-Time Employee Benefits Include: 💰 Competitive Salary & Equity 💹 401(k) Program with a 4% match (US Only) ⚕️ Health, Dental, Vision and Life Insurance 🩼 Short Term and Long Term Disability 🚼 Paid Parental, Medical, Caregiver Leave 🏝 Flexible Time Off (FTO) + Holidays 🚗 Commuter Benefits (In-Office Only) 📱 Monthly Wellness Stipend 🧑‍💻 Autonomous Work Environment 🖥 In Office Set-Up Reimbursement (In-Office Only) 🚀 Quarterly Team Gatherings ☕ In Office Amenities (In-Office Only) Want to learn more about what we are up to? - Meet the Replit Agent - Replit: Make an app for that - Replit Blog - Amjad TED Talk Interviewing + Culture at Replit - Operating Principles - Reasons not to work at Replit To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.

Foster City, CA hybrid FullTime
AWSAzureDocker +3 more

Senior AI Infrastructure Engineer - Training Platform

2mo ago
Scale AI

Scale AI

As a Software Engineer on the Machine Learning Infrastructure team, you will build the "Operating System" for our large-scale GPU clusters. You will architect a high-performance training platform that handles the immense complexity of multi-thousand GPU workloads, ensuring every cycle is used efficiently. Your work directly determines the velocity at which our researchers can train and iterate on the world’s most advanced models. The ideal candidate is a systems expert who thrives on solving the orchestration, networking, and reliability challenges that emerge at massive scale. You will partner closely with researchers to build a seamless, resilient environment that transforms raw compute into breakthrough AI.

$216k - $270k

San Francisco, CA; Seattle, WA; New York, NY onsite
AWSKubernetesPython +8 more

ML Systems Engineer, Robotics

2mo ago
Scale AI

Scale AI

Scale's Physical AI business unit is focused on solving data bottlenecks in Robotics, Autonomous Vehicles, and Computer Vision. This role involves applied research and developing ML pipelines for processing, training, and fine-tuning data collected by Scale, with an emphasis on optimizing algorithms and pipelines for efficient GPU execution in the cloud. You will advance research, shape Scale's offerings, and expand the frontier of data and model evaluation for Physical AI. As an ML Systems Engineer, you will design and build platforms for scalable, reliable, and efficient serving of foundation models tailored for physical agents, powering both internal research and external customer use cases.

$249k - $311k

San Francisco, CA remote
AWSDockerKubernetes +7 more

Senior AI Infrastructure Engineer, Model Serving Platform

2mo ago
Scale AI

Scale AI

As a Software Engineer on the ML Infrastructure team, you will design and build platforms for scalable, reliable, and efficient serving of LLMs. Our platform powers cutting-edge research and production systems, supporting both internal and external use cases across various environments. You'll work in a highly collaborative environment, bridging research and engineering to deliver seamless experiences to our customers and accelerate innovation across the company.

$216k - $270k

San Francisco, CA; New York, NY remote
AWSDockerKubernetes +7 more

Staff Infrastructure Software Engineer, Enterprise AI

2mo ago
Scale AI

Scale AI

Scale GP is seeking a Senior or Staff Infrastructure Engineer to lead the engineering of the 'paved road' for knowledge retrieval and inference engines, defining deployment standards for Agentic workflows at scale. This role bridges complex AI orchestration with world-class infrastructure, ensuring platform reliability for enterprise agents. The ideal candidate is passionate about deep technical work, mentoring, and setting long-term technical strategy while maintaining a hands-on delivery focus. You will architect and implement solutions across multiple cloud providers (GCP, Azure, AWS) for customers in regulated industries.

$252k - $315k

New York, NY; San Francisco, CA remote
AWSAzureKubernetes +13 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.