Senior Software Engineer, Cloud Infrastructure

$200k - $400k San Francisco FullTime

Posted 2mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

Decagon is seeking a Senior Software Engineer to join their Infrastructure team. This role will focus on building and operating the foundational systems that power Decagon's conversational AI platform, including networking, data, ML serving, and developer platforms. You will be responsible for architecting and operating deployments within enterprise customer clouds, ensuring reliability, security, and compliance. This position involves building platforms and abstractions for product teams, managing end-to-end ownership of enterprise deployments, and ensuring the reliability of agentic AI workloads. The role requires a proactive approach to problem-solving in a rapidly evolving technological landscape.

Responsibilities

  • Build and design development and production platforms, including abstractions over cloud infrastructure, Kubernetes, and networking.
  • Scale platforms to accommodate significant growth in usage.
  • Own the end-to-end architecture and lifecycle of deployments in customer-owned cloud environments.
  • Develop runbooks and automation for repeatable deployment processes.
  • Implement monitoring, alerting, and rollback strategies for production systems.
  • Ensure the reliability of AI agent systems, focusing on latency, availability, and graceful degradation.
  • Collaborate with customer platform, security, and DevOps teams, as well as internal Product, Security, Sales, and Customer Success teams.

Requirements

  • 4+ years of experience in core infrastructure, platform engineering, or infrastructure/DevOps.
  • Experience with customer-facing deployment.
  • Deep experience with a major cloud provider (GCP, AWS, or Azure).
  • Extensive experience with Terraform and Kubernetes at scale.
  • Strong understanding of cloud networking fundamentals (VPCs, IAM, DNS, load balancing).
  • Proven track record of operating production systems reliably, including monitoring, on-call, and incident response.
  • Ability to navigate ambiguity and translate stakeholder needs into actionable plans.
  • Clear technical writing skills and experience driving team adoption.
  • Comfort working in a fast-moving environment with rapid change.

Benefits

  • Medical, Dental, and Vision benefits
  • Life Insurance and Disability Benefits
  • Retirement Plan (e.g., 401K, pension)
  • Equity

About decagon

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.