Staff Software Engineer, Data Platform

$231k - $340k Remote New York FullTime

Posted 19d ago

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

Harvey is seeking a Staff Software Engineer to join their central data platform team. This role is crucial for building the foundational systems that enable all teams at Harvey to work with data confidently and independently. You will be responsible for creating frameworks, tooling, and paved paths for product engineers, data engineers, and analysts, focusing on ingestion, warehousing, transformation, and real-time processing. A key aspect of this role involves ensuring data quality, lineage, governance, and handling sensitive data like PII and multi-region residency requirements, which are critical for serving security-conscious institutions.

Responsibilities

  • Own the data platform's architecture and technical direction, treating data infrastructure as a software product.
  • Build and operate the ingestion layer across streaming, batch, CDC, and third-party connectors.
  • Land data into Snowflake with defined freshness, completeness, and cost characteristics.
  • Own the orchestration platform, including scheduling, retries, backfills, and dependency management.
  • Build transformation and compute frameworks for processing data at scale.
  • Design and operate stream processing infrastructure for real-time use cases.
  • Build the trust layer for data quality, observability, lineage, cataloging, and discovery.
  • Build patterns and tooling for PII and sensitive data handling, including classification, masking, retention, and access control.
  • Set the technical bar for data at Harvey through design reviews, standards, documentation, and mentorship.

Requirements

  • 10+ years building and operating production data infrastructure with ownership of dependent systems.
  • Deep experience with cloud data warehouses, preferably Snowflake, including performance tuning and cost management.
  • Hands-on experience building CDC and streaming pipelines with technologies like Kafka, Debezium, Flink, or Spark Streaming.
  • Experience with managed ingestion tooling (Fivetran, Airbyte, or similar).
  • Strong fluency with workflow orchestration at scale (Temporal, Airflow, Dagster, or similar).
  • Strong programming skills in Python and advanced SQL.
  • Experience building frameworks or internal tooling for other engineers.
  • Practical experience with data quality, observability, and lineage.

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.