Staff Software Engineer, Production Engineering

$231k - $340k Remote New York FullTime

Posted 11h ago

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

Harvey is seeking a Production Engineer to build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth.

Responsibilities

  • Design, build, and operate production infrastructure for products and AI workloads.
  • Drive technical direction for compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.
  • Lead complex technical initiatives to improve reliability, scalability, security, and efficiency.
  • Partner with other teams to translate requirements into infrastructure solutions.
  • Establish reusable patterns and tooling for engineering teams.
  • Raise the engineering bar through design reviews, documentation, and mentorship.
  • Build and operate global compute and network infrastructure for high availability and performance.
  • Improve compute utilization and performance for AI workloads.
  • Develop capacity models and fleet lifecycle automation.
  • Operate and improve the Kubernetes platform, including provisioning, networking, and automation.
  • Drive infrastructure cost efficiency through capacity management and optimization.
  • Build secure infrastructure foundations, including IAM, network security, and secrets management.
  • Develop Infrastructure-as-Code and automation frameworks using Terraform or Pulumi.
  • Improve observability, monitoring, alerting, and incident response.
  • Participate in on-call rotation and lead incident response.

Requirements

  • 10+ years of software, infrastructure, site reliability, or production engineering experience.
  • Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or GCP.
  • Strong hands-on experience operating Kubernetes in production.
  • Experience building and operating distributed systems with strong reliability and scalability.
  • Experience with infrastructure automation and Infrastructure-as-Code (Terraform or Pulumi).
  • Strong understanding of compute infrastructure, networking, capacity planning, and fleet management.
  • Experience designing and operating observability systems (monitoring, logging, alerting).
  • Strong understanding of infrastructure security best practices.
  • Track record of driving complex technical initiatives and influencing engineering decisions.
  • Excellent communication skills.
  • Systems-thinking mindset.

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.