Senior Software Engineer - Together Cloud Infrastructure
$160k - $230k • Remote • San Francisco
Posted 1mo ago
About the job
Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. As a Senior AI Infrastructure Engineer, you will play a key role in building the next generation AI cloud platform – a highly available, global, blazing-fast cloud infrastructure that virtualizes cutting-edge ML hardware and enables state-of-the-art ML practitioners with self-serve AI cloud services. This platform serves both our internal SaaS products and our external cloud customers, spanning dozens of data centers across the world.
Responsibilities
- Design, build, and maintain performant, secure, and highly-available backend services/operators for hardware management automation.
- Design and build the IaaS software layer for a new data center with thousands of GPUs.
- Work on a global multi-exabyte high-performance object store for massive datasets.
- Build advanced observability stacks with automated node lifecycle management for distributed pretraining.
- Perform architecture and research for decentralized AI workloads.
- Work on the core, open-source Together AI platform.
- Create services, tools, and developer documentation.
- Create testing frameworks for robustness and fault-tolerance.
Requirements
- 5+ years of professional software development experience.
- Proficiency in at least one backend programming language (Golang desired).
- 5+ years experience writing high-performance, well-tested, production quality code.
- Demonstrated experience building and operating high-performance and/or globally distributed micro-service architectures across cloud providers (AWS, Azure, GCP).
- Excellent communication skills, able to write clear design docs and work effectively with technical and non-technical team members.
- Deep experience with Kubernetes internals (operators, plugins, schedulers) is a big plus.
- Deep experience with VMs/hypervisors (QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV) is a big plus.
- Deep experience with DC networking tech (VLAN, VXLAN, VPN, VPC, OVS/OVN) is a big plus.
- Experience with Cluster API or similar is a big plus.
- Experience with high-performance compute, networking, and/or storage is a big plus.
- Experience virtualizing GPUs and/or Infiniband is a big plus.
- Strong systems knowledge across compute, networking, and storage, including concurrency, memory management, performant I/O, and scale.
- Experience with infrastructure automation tools (Terraform, Ansible), monitoring/observability stacks (Prometheus, Grafana), and CI/CD pipelines (GitHub Actions, ArgoCD).
- Experience building IaaS or PaaS systems at scale is a plus.
- Experience with DPUs/SmartNICs is a plus.
- GPU programming, NCCL, CUDA knowledge is a plus.
Benefits
- Competitive compensation
- Startup equity
- Health insurance
- Other benefits
- Flexibility in terms of remote work