Staff Software Engineer, Production Engineering
$231k - $340k • Remote • New York • FullTime
Posted 11h ago
Job Location
New York
Tech Stack
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
Harvey is seeking a Production Engineer to build and operate its core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. This role is crucial for enabling engineering teams to move quickly and operate reliable services at scale. You will focus on improving the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform, solving complex production challenges across fleet management, capacity planning, automation, and operations. You will collaborate closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth.
Responsibilities
- Design, build, and operate production infrastructure for products and AI workloads.
- Drive technical direction for compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.
- Lead complex technical initiatives to improve reliability, scalability, security, and efficiency.
- Partner with other teams to translate requirements into infrastructure solutions.
- Establish reusable patterns and tooling for engineering teams.
- Raise the engineering bar through design reviews, documentation, and mentorship.
- Build and operate global compute and network infrastructure for high availability and performance.
- Improve compute utilization and performance for AI workloads.
- Develop capacity models and fleet lifecycle automation.
- Operate and improve the Kubernetes platform, including provisioning, networking, and automation.
- Drive infrastructure cost efficiency through capacity management and optimization.
- Build secure infrastructure foundations, including IAM, network security, and secrets management.
- Develop Infrastructure-as-Code and automation frameworks using Terraform or Pulumi.
- Improve observability, monitoring, alerting, and incident response.
- Participate in on-call rotation and lead incident response.
Requirements
- 10+ years of software, infrastructure, site reliability, or production engineering experience.
- Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or GCP.
- Strong hands-on experience operating Kubernetes in production.
- Experience building and operating distributed systems with strong reliability and scalability.
- Experience with infrastructure automation and Infrastructure-as-Code (Terraform or Pulumi).
- Strong understanding of compute infrastructure, networking, capacity planning, and fleet management.
- Experience designing and operating observability systems (monitoring, logging, alerting).
- Strong understanding of infrastructure security best practices.
- Track record of driving complex technical initiatives and influencing engineering decisions.
- Excellent communication skills.
- Systems-thinking mindset.