Software Engineer, Site Reliability

$160k - $350k NYC FullTime

Posted 6mo ago

Job Location

NYC

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Hebbia is seeking a Site Reliability Engineer who approaches the role with a software engineering mindset. You will be responsible for the entire lifecycle of critical production systems, focusing on design, development, and enhancement rather than just operation. This involves writing production-quality code to ensure platform reliability at scale, collaborating with product engineering teams to integrate reliability into architectural decisions from the outset, and building essential internal tooling for all engineers. The role emphasizes coding, including instrumenting services, optimizing performance, developing deployment platforms, and translating incident learnings into permanent architectural improvements.

Responsibilities

  • Own critical production services end-to-end, including design, code review, deployment, operation, and incident response.
  • Profile, benchmark, and rewrite critical code paths to eliminate bottlenecks as Hebbia scales.
  • Lead incident response and foster a post-mortem culture, translating findings into code and architectural improvements.
  • Design and build observability frameworks from scratch, including custom instrumentation, alerting, and debugging tools.
  • Define and enforce SLOs across platform services and establish feedback loops for accountability.
  • Manage capacity planning and cost efficiency, modeling growth, right-sizing infrastructure, and automating resource optimization.
  • Build robust, well-tested internal platforms and deployment tooling.
  • Own and continuously improve CI/CD systems to enable safe and rapid deployments.
  • Embed with product engineering teams as a peer software engineer, contributing to production codebases and co-designing for reliability.
  • Partner on infrastructure security through threat modeling, hardening, and automated compliance tooling.

Requirements

  • 5+ years of software development experience with a track record of writing, shipping, and maintaining production services.
  • Production-grade proficiency in at least one systems or backend language: Go, Python, C++, or Rust.
  • Proven experience as a Production Engineer, SRE, or software engineer with a deep infrastructure focus.
  • Deep understanding of distributed systems.
  • Container orchestration expertise and experience debugging complex distributed failures in production.
  • Working knowledge of OS-level concepts.
  • Cloud platform fluency (AWS preferred).
  • Experience building and maintaining observability stacks.
  • Strong CI/CD pipeline expertise and a track record of improving developer velocity.
  • Experience building platforms for AI/ML workloads or high-throughput document processing pipelines is a plus.

Benefits

  • Unlimited PTO
  • Medical, Dental, and Vision insurance
  • 401K
  • Catered lunch daily
  • Doordash dinner credit
  • 3 months parental leave for non-birthing parent
  • 4 months parental leave for birthing parent
  • $15k lifetime fertility benefit
  • Competitive new hire equity grant with unmatched upside potential

About Hebbia

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.