Senior Product Operations Manager, Evaluation

Remote San Francisco FullTime

Posted 1mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Harvey is transforming legal and professional services by integrating frontier agentic AI, an enterprise-grade platform, and deep domain expertise. We are building a generational company at an inflection point, with strong product-market fit and significant investor support, scaling rapidly to define a new category. Our team is driven by decisiveness, simplicity, and a commitment to continuous improvement, operating with intensity and a focus on customer needs. This is an opportunity to do your best work alongside ambitious colleagues.

We are seeking a technical, systems-minded operator to build and scale the evaluation engine for Harvey's platform. As we expand globally, ensuring our models are reliable, accurate, and jurisdictionally correct is critical, and evaluation complexity is increasing significantly. You will work closely with Applied Legal Researchers, Product, Engineering, AI Research, and human data providers to operationalize evaluation methodologies and integrate them into our product development lifecycle. This role involves creating the workflows, systems, and tooling to establish evaluation as a core product capability, offering high ownership for individuals who thrive in ambiguity and enjoy building structure within a global AI company.

Responsibilities

  • Build and scale systems for model and product evaluations across Harvey.
  • Manage the evaluation request queue, including intake, triage, and prioritization, to address the highest-value coverage gaps.
  • Integrate evaluation workflows and readiness checkpoints into the product development lifecycle.
  • Establish a single source of truth for evaluation status, results, history, and launch readiness.
  • Translate expert-designed evaluation methodologies into scalable, repeatable operational processes.
  • Manage human data providers and establish an internal contract-attorney pipeline to ensure evaluation quality meets legal standards.
  • Collaborate with Engineering and Research to enhance evaluation tooling, automation, and dashboards.
  • Drive evaluation readiness for major product and model launches across various geographies and jurisdictions.
  • Document and operationalize evaluation governance as complexity increases.
  • Contribute to defining how Harvey ensures model accuracy, reliability, and trust at a global scale.

Requirements

  • 4-7+ years in technical program management, product operations, research operations, or evaluation/benchmarking roles.
  • Experience with ML/AI evaluations, benchmarking frameworks, or scientific workflows.
  • Proficiency with statistical methodologies and tools like SQL or Python for data interpretation.
  • Strong business acumen with an ROI-focused mindset for scaling.
  • Ability to work effectively with legal experts and operationalize complex evaluation methodologies.
  • Excellent cross-functional coordination skills across Product, Engineering, Research, and data providers/vendors.
  • High attention to detail with a focus on clarity, rigor, and reproducibility.
  • Ability to navigate evolving landscapes and bring order to complex systems.
  • Strong communication skills, capable of translating technical nuances for diverse stakeholders.
  • Willingness to perform all necessary tasks to ensure the success of evaluation systems, including documentation and pipeline issue diagnosis.

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.