Software Engineer - GPU Networking & Distributed Systems

Remote San Francisco FullTime

Posted 6mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Baseten is building the global operating system for distributed, heterogeneous AI hardware, powering mission-critical inference for leading AI companies. As LLM and multi-modal workloads scale, the network becomes a critical component. This role focuses on leading GPU Networking efforts, making RDMA a first-class building block and optimizing distributed inference. You will architect the software fabric that unifies thousands of GPUs, co-optimizing communication and computation for advanced AI workloads.

Responsibilities

  • Integrate RDMA/RoCE/InfiniBand capabilities into the inference stack to improve bandwidth and latency.
  • Implement and tune networking layers for efficient Disaggregated KV Cache Offload and WideEP.
  • Enable sub-10-second startup times for large models by working with checkpointing and storage mechanisms.
  • Characterize and validate networking performance on cutting-edge GPU hardware.
  • Design tools for visualizing packet flow, congestion, and bandwidth across GPU interconnects.
  • Optimize communication libraries (NCCL, NVSHMEM) and potentially write custom communication kernels.

Requirements

  • Deep experience with high-performance networking protocols (InfiniBand, RoCE v2).
  • Fluency in C++ or Python, with the ability to bridge high-level logic and hardware.
  • Deep understanding of memory hierarchy in modern NVIDIA architectures (H100/Blackwell) and optimization techniques.
  • Willingness to dive deep into source code (e.g., TensorRT-LLM) and debug complex issues.
  • Ability to discern when to use off-the-shelf solutions versus building custom ones.

Benefits

  • Competitive compensation and meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents.
  • Flexible PTO policy including company-wide Winter Break.
  • Paid parental leave.
  • Fertility and family-building stipend.
  • Company-facilitated 401(k).
  • Exposure to various ML startups for learning and networking.

About Baseten

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.