Senior Software Engineer, Production Engineering

$161k - $242k Remote San Francisco FullTime

Posted 11h ago

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

Harvey is seeking a Production Engineer to help build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth. The ideal candidate will have a systems-thinking mindset and a passion for building simple, reliable, and scalable systems.

Responsibilities

  • Design, build, and operate production infrastructure for products and AI workloads.
  • Drive technical direction for compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.
  • Lead cross-functional initiatives to improve reliability, scalability, security, operational efficiency, and infrastructure cost.
  • Partner with other engineering teams to translate requirements into infrastructure solutions.
  • Establish reusable patterns, tooling, and paved paths for engineering teams.
  • Raise the engineering bar through design reviews, documentation, operational rigor, and mentorship.
  • Build and operate global compute and network infrastructure for high availability, scalability, reliability, and performance.
  • Improve compute utilization, performance, and service availability for AI workloads.
  • Develop capacity models, demand forecasts, and fleet lifecycle automation.
  • Operate and improve the Kubernetes platform, including cluster management, networking, monitoring, and automation.
  • Drive infrastructure cost efficiency through capacity management and resource optimization.
  • Build secure infrastructure foundations, including IAM, network security, secrets management, and compliance.
  • Develop scalable Infrastructure-as-Code and automation frameworks using tools like Terraform or Pulumi.
  • Improve observability, monitoring, alerting, incident response, and operational readiness.
  • Participate in on-call rotation, lead incident response, and implement engineering improvements based on production learnings.

Requirements

  • 5+ years of software, infrastructure, site reliability, or production engineering experience.
  • Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform.
  • Strong hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, and reliability.
  • Experience building and operating distributed systems with strong reliability, scalability, and performance characteristics.
  • Experience with infrastructure automation and Infrastructure-as-Code using tools such as Terraform or Pulumi.
  • Strong understanding of compute infrastructure, networking, capacity planning, fleet management, and production operations.
  • Experience designing and operating observability systems, including monitoring, logging, alerting, and incident response.
  • Strong understanding of infrastructure security, including IAM, network security, secrets management, and compliance best practices.
  • Track record of driving complex, cross-functional technical initiatives and influencing engineering decisions.
  • Excellent communication skills and ability to explain technical concepts clearly.
  • Systems-thinking mindset and passion for building simple, reliable, and scalable systems.

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.