Software Engineer - GPU Networking & Distributed Systems
Remote • San Francisco • FullTime
Posted 6mo ago
About the job
Baseten is building the global operating system for distributed, heterogeneous AI hardware, powering mission-critical inference for leading AI companies. As LLM and multi-modal workloads scale, the network becomes a critical component. This role focuses on leading GPU Networking efforts, making RDMA a first-class building block and optimizing distributed inference. You will architect the software fabric that unifies thousands of GPUs, co-optimizing communication and computation for advanced AI workloads.
Responsibilities
- Integrate RDMA/RoCE/InfiniBand capabilities into the inference stack to improve bandwidth and latency.
- Implement and tune networking layers for efficient Disaggregated KV Cache Offload and WideEP.
- Enable sub-10-second startup times for large models by working with checkpointing and storage mechanisms.
- Characterize and validate networking performance on cutting-edge GPU hardware.
- Design tools for visualizing packet flow, congestion, and bandwidth across GPU interconnects.
- Optimize communication libraries (NCCL, NVSHMEM) and potentially write custom communication kernels.
Requirements
- Deep experience with high-performance networking protocols (InfiniBand, RoCE v2).
- Fluency in C++ or Python, with the ability to bridge high-level logic and hardware.
- Deep understanding of memory hierarchy in modern NVIDIA architectures (H100/Blackwell) and optimization techniques.
- Willingness to dive deep into source code (e.g., TensorRT-LLM) and debug complex issues.
- Ability to discern when to use off-the-shelf solutions versus building custom ones.
Benefits
- Competitive compensation and meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employee and dependents.
- Flexible PTO policy including company-wide Winter Break.
- Paid parental leave.
- Fertility and family-building stipend.
- Company-facilitated 401(k).
- Exposure to various ML startups for learning and networking.