Staff Software Engineer, Model Infrastructure

$236k - $290k San Francisco FullTime

Posted 8d ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

Harvey is transforming professional services by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. We are a fast-growing company with strong product-market fit, seeking individuals who want to do the best work of their careers. As a Staff Software Engineer on the Model Infrastructure team, you will lead the design and development of the systems that power every AI request at Harvey. You will partner closely with AI Research, Product Engineering, and Infrastructure teams to build a highly reliable, scalable, observable, and efficient platform.

Responsibilities

  • Lead the design and implementation of Harvey's Model Infrastructure platform.
  • Build systems for high availability, low latency, and operational excellence in AI inference.
  • Design and improve the Unified Model Controller (UMC) and Model Selector platform for automated model degradation detection and intelligent traffic routing.
  • Develop systems for model provisioning, capacity management, failover, and traffic engineering across multiple AI providers.
  • Integrate new model providers and maintain provider APIs and SDKs to enable rapid adoption of emerging frontier models.
  • Improve observability through health dashboards, alerting, token usage analytics, cost reporting, and end-to-end telemetry.
  • Partner with Product Engineering to support model launches, experimentation, and proactive monitoring of production AI workloads.
  • Drive infrastructure efficiency through capacity planning, utilization optimization, and cost visibility.
  • Collaborate with AI Research to build the infrastructure foundation for future model evaluation, training, and deployment.
  • Lead cross-functional technical initiatives and mentor engineers.

Requirements

  • 7+ years of software engineering experience building large-scale distributed systems.
  • Experience designing and operating highly available production services.
  • Strong programming skills in Go, Java, Python, Rust, or C++.
  • Deep understanding of distributed systems, cloud infrastructure, networking, and observability.
  • Experience leading technical projects across multiple engineering teams.
  • Ability to balance long-term architecture with pragmatic execution.
  • Strong communication and collaboration skills.
  • Passion for building foundational platforms that enable other engineering teams.
  • Experience with AI infrastructure, LLM serving, or machine learning platforms (nice to have).
  • Experience with model routing, inference gateways, or policy-based serving systems (nice to have).
  • Experience working with OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, or open-source LLMs (nice to have).
  • Experience with Kubernetes, cloud infrastructure, and service mesh technologies (nice to have).
  • Experience with large-scale observability and SRE best practices (nice to have).
  • Experience with data infrastructure technologies such as Kafka, Spark, Flink, Airflow, or Iceberg (nice to have).
  • Familiarity with GPU infrastructure or model training platforms (nice to have).

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.