Software Engineer, Data Foundations

$180k - $300k San Francisco, CA

Posted 1mo ago

Job Location

San Francisco, CA

Tech Stack

Remote Work Policy

On-site

Categories

Applied AI Engineer

About the job

Glean is seeking a Software Engineer to join its Data Foundations team. This team is responsible for the entire data ingestion and management layer that powers Glean's Search, AI Assistant, and Agent products. The work directly impacts the quality, freshness, and trustworthiness of the knowledge users interact with daily. You will contribute to building and scaling connectors to various SaaS and on-prem systems, handling complex data synchronization, and transforming raw enterprise content into structured representations optimized for search and LLM reasoning. This role involves designing data schemas, enrichment pipelines, and integrating AI product capabilities to automate tasks and enhance data utilization.

Responsibilities

  • Build and scale connectors to various SaaS and on-prem systems (e.g., Google Workspace, Microsoft 365, Slack, Salesforce, Jira, ServiceNow, GitHub).
  • Handle full syncs, low-latency incremental updates via webhooks/APIs, rate-limiting, and complex authentication flows.
  • Develop advanced datasource capabilities such as actions, live-fetch, and query language support.
  • Transform unstructured enterprise content into rich, structured, permission-aware representations for search and LLM reasoning.
  • Design document schemas and enrichment pipelines, including entity extraction and access-graph propagation.
  • Expand AI product capabilities through deep integrations for task automation, complex data-grounded queries, and live data enhancement.
  • Ensure end-to-end correctness, freshness, and performance for petabyte-scale data flows.
  • Solve complex distributed systems problems related to ordering, idempotency, exactly-once processing, backpressure, and retries.

Requirements

  • Experience in building and scaling data ingestion and management layers.
  • Proficiency in handling data synchronization, including full syncs and incremental updates.
  • Experience with APIs, webhooks, rate-limiting, and authentication flows.
  • Familiarity with transforming unstructured data into structured formats.
  • Experience in designing data schemas and enrichment pipelines.
  • Understanding of distributed systems concepts like ordering, idempotency, and backpressure.
  • Ability to work with petabyte-scale data flows.

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.