Senior Software Engineer, Core Infrastructure
$200k - $250k • Remote • San Francisco • FullTime
Posted 6mo ago
About the job
As a Software Engineer on the Core Infrastructure team at Harvey, you will play a critical role in designing and building new infrastructure systems while equally scaling and strengthening our existing infrastructure. This foundation powers every user interaction with Harvey, processing billions of prompt tokens and millions of daily requests across our global legal AI platform. You will work in an environment balanced between innovation and operational excellence, ensuring Harvey remains resilient and efficient as it scales products, regions, customers, and usage. Your contributions will directly impact the reliability, scalability, and security of our platform as we serve the world's leading law firms and professional service providers.
Responsibilities
- Design and build scalable, fault-tolerant infrastructure systems for Harvey's AI platform across multiple cloud regions.
- Own and evolve multi-cloud infrastructure (Azure, GCP), including Kubernetes orchestration, networking, and container management.
- Lead technical initiatives around observability, incident response, and operational excellence.
- Architect and optimize distributed systems for reliability, including load balancing, quota management, and failover mechanisms.
- Partner with Product Engineering and Security teams to ensure infrastructure accelerates product development.
- Drive infrastructure-as-code practices using tools like Terraform and Pulumi for reproducible deployments.
- Mentor junior engineers and raise the technical bar through code and design reviews.
Requirements
- 4+ years of experience in Infrastructure Engineering or Platform Engineering in a production environment.
- Proven track record building and scaling complex, large-scale distributed systems.
- Deep proficiency with cloud infrastructure platforms (Azure preferred; GCP or AWS experience is transferable).
- Strong fluency in Infrastructure as Code (IaC) tools such as Terraform, Pulumi, or CloudFormation.
- Solid understanding of Kubernetes, container orchestration, networking, and cloud security at scale.
- Experience with observability tools (e.g., Datadog, Sentry) and incident response practices (e.g., PagerDuty, Incident.io).
- Strong programming skills in Python, Go, or similar languages.
- Excellent problem-solving skills and a commitment to operational excellence.