Senior Software Engineer Together Cloud Infrastructure

Amsterdam

Posted 1mo ago

Remote Work Policy

On-site

Categories

AI Infrastructure Engineer

About the job

Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. As a Senior AI Infrastructure Engineer, you will play a key role in building the next generation AI cloud platform – a highly available, global, blazing-fast cloud infrastructure that virtualizes cutting-edge ML hardware and enables state-of-the-art ML practitioners with self-serve AI cloud services. This platform serves both our internal SaaS products and our external cloud customers, spanning dozens of data centers across the world.

Responsibilities

  • Design, build, and maintain performant, secure, and highly-available backend services/operators for hardware management automation.
  • Design and build the IaaS software layer for new data centers with thousands of GPUs.
  • Work on a global multi-exabyte high-performance object store for massive datasets.
  • Build advanced observability stacks with automated node lifecycle management.
  • Perform architecture and research for decentralized AI workloads.
  • Work on the core, open-source Together AI platform.
  • Create services, tools, and developer documentation.
  • Create testing frameworks for robustness and fault-tolerance.

Requirements

  • 5+ years of professional software development experience.
  • Proficiency in at least one backend programming language (Golang desired).
  • 5+ years experience writing high-performance, well-tested, production quality code.
  • Demonstrated experience building and operating high-performance and/or globally distributed micro-service architectures.
  • Excellent communication skills.
  • Deep experience with Kubernetes internals (operators, plugins, schedulers, patches).
  • Deep experience with VMs/hypervisors (QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV).
  • Deep experience with DC networking tech (VLAN, VXLAN, VPN, VPC, OVS/OVN).
  • Experience with Cluster API or similar.
  • Experience working on high-performance compute, networking, and/or storage.
  • Experience virtualizing GPUs and/or Infiniband.
  • Strong systems knowledge across compute, networking, and storage.
  • Experience with infrastructure automation tools (Terraform, Ansible).
  • Experience with monitoring/observability stacks (Prometheus, Grafana).
  • Experience with CI/CD pipelines (GitHub Actions, ArgoCD).
  • Experience building IaaS or PaaS systems at scale.
  • Experience with DPUs/SmartNICs.
  • GPU programming, NCCL, CUDA knowledge.

About Together AI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.