Engineering Manager, Cloud Monitoring Services Platform

$215k - $260k • San Francisco, CA - US • FullTime

Posted 3h ago

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Machine Learning Engineer

About the job

Crusoe is building cloud infrastructure for AI workloads and is seeking an Engineering Manager to lead the Platform team within Cloud Monitoring Services. This team is responsible for observability across Crusoe Cloud, including metrics, logs, alerting, and the telemetry agent. The Platform team specifically manages the time series and log storage systems, as well as the query layer that serves all dashboards and API calls. This is a first-line management role focused on people leadership and technical guidance, with success measured by the team's accomplishments in areas like query latency, retention, and storage cost.

Responsibilities

  • Grow and develop a team of 4-6 engineers, managing 1:1s, career growth, performance, and team health.
  • Own the time series and log storage systems and the query layer, including performance under load and cost.
  • Plan and sequence work across a roadmap balancing customer-facing features and infrastructure improvements.
  • Set the technical direction for the storage and query platform in partnership with Staff engineers.
  • Maintain a high operational bar for availability, query performance, and correctness.
  • Own the team's on-call rotation and the operational health of the storage and query stack.
  • Manage cost and scale through retention policies, downsampling, cardinality, and storage tiering.
  • Coordinate dependencies and escalate issues across team boundaries with collection, ingestion, and platform teams.
  • Shape the team by working with recruiting, running interviews, and onboarding new hires.
  • Collaborate with product, neighboring infrastructure teams, and leadership to align priorities.

Requirements

  • 5+ years of hands-on engineering experience, ideally in backend or distributed systems (databases, storage engines, query engines, streaming pipelines).
  • Hands-on familiarity with Kubernetes.
  • 2+ years of direct software engineering management experience, including performance cycles and career conversations.
  • Experience managing a team through delivery with an active roadmap and live customer commitments.
  • Experience coaching senior and staff-level engineers and partnering with technical leads.
  • Proven track record of delivering multi-phase projects against deadlines, balancing infrastructure and feature work.
  • Experience running a team that owns a customer-dependent system during incidents.
  • Strong people management skills, including giving direct and kind feedback.
  • A desire for engineers around them to grow and measuring success through team accomplishments.

About crusoe

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.