Member of Technical Staff - Platform Engineering

New York FullTime

Posted 7mo ago

Job Location

New York

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

AI needs a new infrastructure layer, and Modal is building it. We provide instant GPU access, sub-second container starts, and native storage, enabling customers to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. As a rapidly growing cloud infrastructure company, we are seeking to dramatically improve our platform's reliability while scaling our team and customer base. This role is ideal for individuals with deep systems thinking, a passion for reliability, and a drive to enable others to move faster at scale.

Responsibilities

  • Identify architectural changes to improve reliability and performance.
  • Foster a culture of reliability across the engineering organization.
  • Define and implement operational processes such as deployments and upgrades.
  • Operate systems like Kubernetes, Postgres, and Redis.
  • Participate in on-call rotations and respond to production incidents.

Requirements

  • 5+ years of experience writing high-quality production code.
  • 2+ years of on-call experience for critical production services.
  • Strong cloud skills with deep familiarity in at least one hyperscaler cloud (AWS preferred).
  • Familiarity with auto-scaling, fleet management, and capacity planning at scale.
  • Experience operating databases, monitoring, CI/CD, and other infrastructure at scale.
  • Experience owning and scaling Kubernetes clusters to thousands of nodes is a plus.
  • Experience with systems safety research (e.g., STAMP) and control theory is a plus.

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.