Staff Software Engineer, Core Infrastructure
$201k - $264k • New York • FullTime
Posted 3mo ago
About the job
Harvey is transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. This is a rare chance to help build a generational company at a true inflection point, scaling fast and defining a new category. The work is ambitious, the bar is high, and the opportunity for growth is unmatched. As a Staff Software Engineer on the Core Infrastructure team, you will play a critical role in designing and building new infrastructure systems while scaling and strengthening existing ones. Our infrastructure powers every user interaction with Harvey, processing billions of prompt tokens and millions of daily requests across our global legal AI platform. You'll work in an environment balanced between innovation and operational excellence, ensuring Harvey remains resilient and efficient as it scales products, regions, customers, and usage. Your contributions will directly impact the reliability, scalability, and security of our platform.
Responsibilities
- Design and build scalable, fault-tolerant infrastructure systems for Harvey's AI platform across multiple cloud regions.
- Own and evolve multi-cloud infrastructure (Azure, GCP), including Kubernetes orchestration, networking, and container management.
- Lead technical initiatives for observability, incident response, and operational excellence.
- Architect and optimize distributed systems for reliability, including load balancing, quota management, and failover mechanisms.
- Partner with Product Engineering and Security teams to ensure infrastructure accelerates product development.
- Drive infrastructure-as-code practices using tools like Terraform and Pulumi for reproducible deployments.
- Mentor engineers and raise the technical bar through code and design reviews.
Requirements
- 10+ years of experience in Infrastructure Engineering or Platform Engineering in a production environment.
- Long track record building and scaling complex, large-scale distributed systems.
- Deep proficiency with cloud infrastructure platforms (Azure preferred; GCP or AWS experience transfers well).
- Strong fluency with Infrastructure as Code (IaC) tools like Terraform, Pulumi, or CloudFormation.
- Solid understanding of Kubernetes, container orchestration, networking, and cloud security at scale.
- Experience with observability tools (Datadog, Sentry) and incident response practices (PagerDuty, Incident.io).
- Strong programming skills in Python, Go, or similar languages.
- Excellent problem-solving skills and a commitment to operational excellence.