Staff+ Software Engineer, Observability
$320k - $485k • San Francisco, CA | New York City, NY | Seattle, WA
Posted 8h ago
Job Location
San Francisco, CA | New York City, NY | Seattle, WA
Tech Stack
Remote Work Policy
On-site
Categories
AI Infrastructure Engineer
About the job
Anthropic is seeking Software Engineers to join its Observability team within the Infrastructure organization. This team is responsible for the monitoring and telemetry infrastructure that all engineers and researchers at Anthropic rely on, encompassing metrics, logging pipelines, distributed tracing, profiling, error analytics, alerting, and actionable dashboards. By joining this team, you will directly contribute to the reliability and operational excellence of Anthropic's research and product systems. As Anthropic's infrastructure scales across large GPU, TPU, and Trainium clusters, the volume and complexity of operational data are increasing significantly, with many challenging problems residing below the application layer. The team is building next-generation observability systems, including high-throughput telemetry pipelines, fleet-wide continuous profiling, eBPF-based tracing and network visibility, and agentic diagnostic tools, to enable engineers to detect and resolve issues rapidly.
Responsibilities
- Design and build scalable telemetry ingest and storage pipelines for metrics, logs, traces, and error data across Anthropic's multi-cluster infrastructure.
- Build observability solutions providing deep, low-overhead visibility into system behavior across the fleet.
- Own and evolve core observability platforms, leading migrations and architectural improvements for reliability, cost reduction, and scalability.
- Develop instrumentation libraries, SDKs, and eBPF-based auto-instrumentation for high-quality telemetry emission, with or without code changes.
- Reduce mean time to detection and resolution by building cross-signal correlation, unified query interfaces, and AI-assisted diagnostic tooling.
- Improve fleet-wide efficiency by translating continuous profiling and utilization telemetry into actionable optimization insights for CPU, memory, and accelerator fleets.
- Partner with Research, Inference, Product, and Infrastructure teams to ensure observability solutions meet their specific needs.
Requirements
- Experience building and operating large-scale observability or monitoring infrastructure.
- Hands-on experience with observability signals end-to-end, from instrumentation through ingest to query and analysis.
- Understanding of high-throughput telemetry pipelines and the tradeoffs in collecting, storing, and querying operational data at scale.
- Comfort debugging below the application layer, in the kernel, network stack, or hardware.
- Excellent communication skills and ability to partner with internal teams to improve operational visibility and incident response.
- Ability to work independently on ambiguous, high-impact technical challenges.
- Bachelor’s degree or equivalent combination of education, training, and/or experience.
- A field of study relevant to the role as demonstrated through coursework, training, or professional experience.