Software Engineer, Data Foundations
$180k - $300k • San Francisco, CA
Posted 1mo ago
About the job
Glean is seeking a Software Engineer to join its Data Foundations team. This team is responsible for the entire data ingestion and management layer that powers Glean's Search, AI Assistant, and Agent products. The work directly impacts the quality, freshness, and trustworthiness of the knowledge users interact with daily. You will contribute to building and scaling connectors to various SaaS and on-prem systems, handling complex data synchronization, and transforming raw enterprise content into structured representations optimized for search and LLM reasoning. This role involves designing data schemas, enrichment pipelines, and integrating AI product capabilities to automate tasks and enhance data utilization.
Responsibilities
- Build and scale connectors to various SaaS and on-prem systems (e.g., Google Workspace, Microsoft 365, Slack, Salesforce, Jira, ServiceNow, GitHub).
- Handle full syncs, low-latency incremental updates via webhooks/APIs, rate-limiting, and complex authentication flows.
- Develop advanced datasource capabilities such as actions, live-fetch, and query language support.
- Transform unstructured enterprise content into rich, structured, permission-aware representations for search and LLM reasoning.
- Design document schemas and enrichment pipelines, including entity extraction and access-graph propagation.
- Expand AI product capabilities through deep integrations for task automation, complex data-grounded queries, and live data enhancement.
- Ensure end-to-end correctness, freshness, and performance for petabyte-scale data flows.
- Solve complex distributed systems problems related to ordering, idempotency, exactly-once processing, backpressure, and retries.
Requirements
- Experience in building and scaling data ingestion and management layers.
- Proficiency in handling data synchronization, including full syncs and incremental updates.
- Experience with APIs, webhooks, rate-limiting, and authentication flows.
- Familiarity with transforming unstructured data into structured formats.
- Experience in designing data schemas and enrichment pipelines.
- Understanding of distributed systems concepts like ordering, idempotency, and backpressure.
- Ability to work with petabyte-scale data flows.