Engineering Manager, Production Engineering
$260k - $340k • Remote • San Francisco • FullTime
Posted 17d ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Harvey is seeking a Senior Engineering Manager to lead the Infrastructure Foundation & Production Quality Engineering organization. This team is crucial for building and operating Harvey's core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. The role involves owning the reliability, scalability, security, and efficiency of the infrastructure platform, leading a team of engineers, and partnering with various departments to ensure the infrastructure scales with rapid growth. This is a leadership role reporting to the Head of Infrastructure, focused on shaping the future of Harvey's infrastructure platform.
Responsibilities
- Lead, mentor, and grow a team of high-performing infrastructure engineers.
- Foster a culture of operational excellence, engineering quality, customer ownership, and continuous improvement.
- Partner with Engineering, Security, Product, and AI Infrastructure leaders to define long-term infrastructure strategy and execution priorities.
- Drive technical direction for compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.
- Lead cross-functional initiatives to improve reliability, scalability, security, operational efficiency, and infrastructure cost optimization.
- Own and operate Harvey's global compute and network infrastructure, ensuring high availability, scalability, reliability, and performance.
- Manage compute resources to maximize utilization, performance, and service availability while supporting rapidly growing AI workloads.
- Lead capacity planning, demand forecasting, and fleet lifecycle management.
- Operate and continuously improve Harvey's Kubernetes platform.
- Own Harvey's Temporal-based workflow orchestration platform.
- Drive infrastructure cost optimization through capacity management, resource rightsizing, workload efficiency improvements, and utilization monitoring.
- Build and maintain secure infrastructure foundations, including identity and access management, network isolation, secrets management, auditing, and compliance controls.
- Develop scalable Infrastructure-as-Code and automation frameworks using technologies such as Terraform and Pulumi.
- Establish comprehensive observability, monitoring, alerting, incident response, and operational readiness practices.
Requirements
- 7+ years of software or infrastructure engineering experience.
- 5+ years leading engineering teams.
- Deep expertise operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform.
- Strong hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, and reliability.
- Experience building and operating large-scale distributed systems with strong reliability, scalability, and performance characteristics.
- Experience with infrastructure automation and Infrastructure-as-Code using tools such as Terraform or Pulumi.
- Strong understanding of compute infrastructure, networking, capacity planning, fleet management, and production operations.
- Experience designing and operating observability platforms, including monitoring, logging, alerting, and incident response processes.
- Strong understanding of infrastructure security, including IAM, network security, secrets management, and compliance best practices.
- Demonstrated success leading complex cross-functional technical initiatives and influencing engineering strategy across organizations.
- Excellent communication skills.