Senior Product Operations Manager, Evaluation Quality
$155k - $233k • Remote • San Francisco • FullTime
Posted 1d ago
Job Location
San Francisco
Tech Stack
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Harvey is transforming legal and professional services by integrating frontier agentic AI with an enterprise-grade platform and deep domain expertise. This role is a unique opportunity to contribute to a rapidly scaling, category-defining company. We are seeking a senior operator to take ownership of the quality bar for Harvey's human evaluations, ensuring that as evaluation volume grows, the output remains trusted and decision-grade. You will partner closely with Evaluation Operations, Applied Legal Researchers, Product, Engineering, and AI Research teams to set standards for evaluation data and methodology, and certify results before they are used for product shipping.
Responsibilities
- Define and own the quality bar for human evaluations, ensuring outputs are decision-grade for Product, Engineering, and AI Research.
- Author and maintain evaluation guidelines, instructions, benchmarks, and gold references.
- Serve as the arbiter for ambiguous or disputed judgments.
- Standardize and streamline rubric and evaluation design into repeatable templates and documented methodology.
- Own contract-attorney quality, including onboarding, calibration training, inter-rater reliability, and feedback loops.
- Define and maintain a failure-mode/error taxonomy for structured, prioritized signal analysis.
- Run QA on vendor and contract-attorney deliverables before results inform launch decisions.
- Ensure quality bar consistency across jurisdictions and non-English geographies.
- Establish a standard, documented method for analyzing evaluation results and build operational dashboards.
- Support Applied Legal Research (ALR) in certifying evaluation soundness before scaling.
Requirements
- 6+ years in product operations, research operations, evaluation/QA operations, or quality program management.
- Proven track record of owning quality inputs for complex, expert-driven or human-in-the-loop work.
- Experience onboarding, training, and calibrating distributed expert raters and managing feedback loops.
- Proficiency in measurement concepts like calibration, inter-rater reliability, sampling, and rubric design.
- Comfort interpreting evaluation data, with or without AI tool support.
- Experience scaling and streamlining quality processes under pressure, emphasizing documentation and reproducibility.
- Ability to work with domain experts and translate nuanced judgment into repeatable standards.
- Strong cross-functional coordination skills across Product, Engineering, Research, ALR, and vendors.
- Excellent communication skills to build credibility and trust with stakeholders.
- Demonstrated bias to action and high ownership.