Member of Technical Staff - Lead, Machines

New York FullTime

Posted 11h ago

Job Location

New York

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Modal is building the next-generation infrastructure layer for AI, enabling customers to instantly access GPUs, achieve sub-second container starts, and serve low-latency inference and model fine-tuning at scale. We are seeking a strong technical lead to guide our engineers in designing, building, and maintaining the high-performance systems that power our serverless platform. This role will focus on Modal's machines layer, encompassing the fleet of bare metal and cloud hosts, and the control plane responsible for their provisioning, imaging, monitoring, and repair. You will own the entire machine lifecycle, from hardware acceptance and benchmarking to network bring-up, kernel and image management, and automated remediation of unhealthy hosts. This is a hands-on leadership position where you will manage a team of 3-8 engineers while contributing across the stack, shaping our long-term technical direction.

Responsibilities

  • Guide engineers in designing, building, and maintaining the serverless platform's systems.
  • Lead the team responsible for Modal's machines layer, including bare metal and cloud hosts.
  • Own the full lifecycle of a machine, from hardware acceptance and benchmarking to network bring-up and kernel/image management.
  • Manage the provisioning, imaging, monitoring, and repair of the machine fleet.
  • Track and automate the remediation of unhealthy hosts.
  • Stay hands-on across the stack, including BMCs, firmware, PXE, bootloaders, Linux networking, drivers, and distributed control-plane services.
  • Shape the long-term technical path for the machines layer.
  • Contribute to automatic remediation of unhealthy machines to maximize uptime.
  • Oversee automatic integration of new servers into the fleet.
  • Improve network health monitoring and reliability across datacenters.
  • Standardize bare metal network configuration.
  • Develop automatic hardware acceptance testing and benchmarking processes.
  • Manage the custom network bootloader, machine image pipeline, and kernel/firmware management.

Requirements

  • 7+ years of experience writing high-quality production code.
  • 3+ years of direct people management experience, including project planning, growth, and performance conversations.
  • Experience operating large fleets of physical hardware at scale (bare metal provisioning, BMC/IPMI, PXE, network boot, firmware) or building their control planes.
  • Strong cloud skills.
  • Strong knowledge of low-level operating system foundations (Linux kernel, drivers, networking, file systems, containers).
  • Experience working with hardware and colocation providers, including acceptance testing and benchmarking.
  • Track record of setting technical direction and driving architectural decisions.
  • Willingness to participate in on-call rotations and respond to production incidents.

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.