Senior Software Engineer, Data Platform

$193k - $290k Remote New York FullTime

Posted 19d ago

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

Harvey is seeking a Senior Software Engineer to join our central data platform team. This role is crucial for building the systems that empower all teams at Harvey to work with data confidently and independently. You will be responsible for creating frameworks, tooling, and paved paths for product engineers, data engineers, and analysts, focusing on their leverage and trust in our data systems. The initial focus will be on establishing a reliable foundation for ingestion and warehousing, including streaming and batch processing into Snowflake, CDC, orchestration, and schema evolution, while adhering to strict data sensitivity requirements. The role will expand to encompass transformation and compute frameworks, self-serve tooling, real-time stream processing, and robust quality, lineage, and governance layers, all while ensuring PII handling and multi-region data residency are satisfied by design.

Responsibilities

  • Own the data platform's architecture and technical direction, treating data infrastructure as a software product.
  • Build and operate the ingestion layer across streaming, batch, CDC, and third-party connectors.
  • Land data into Snowflake with defined freshness, completeness, and cost characteristics.
  • Own the orchestration platform, including scheduling, retries, backfills, and dependency management.
  • Build transformation and compute frameworks and self-serve tooling for data processing.
  • Design and operate stream processing infrastructure for real-time use cases.
  • Build the trust layer, including quality, observability, lineage, cataloging, and discovery.
  • Build patterns and tooling for PII, sensitive data, and multi-region residency requirements.
  • Set the technical bar for data at Harvey through design reviews, standards, documentation, and mentorship.

Requirements

  • 5+ years building and operating production data infrastructure with ownership of dependent systems.
  • Deep experience with cloud data warehouses, preferably Snowflake (BigQuery, Databricks, or Redshift experience is transferable).
  • Hands-on experience building CDC and streaming pipelines with technologies like Kafka, Debezium, Flink, or Spark Streaming.
  • Experience with managed ingestion tooling (Fivetran, Airbyte, or similar).
  • Strong fluency with workflow orchestration at scale (Temporal, Airflow, Dagster, or similar).
  • Strong programming skills in Python and advanced SQL.
  • Experience building frameworks or internal tooling for other engineers.
  • Practical experience with data quality, observability, and lineage.

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.