Software Engineer, GPU Infrastructure (HPC)

Remote Canada FullTime

Posted 6mo ago

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

Cohere is seeking a Staff Software Engineer to join our internal infrastructure team, responsible for building and operating world-class infrastructure and tools for training, evaluating, and serving Cohere's foundational AI models. You will work closely with AI researchers to support their AI workload needs on cutting-edge systems, focusing on stability, scalability, and observability. This role involves building and operating superclusters across multiple clouds, directly accelerating the development of industry-leading AI models. Participation in a 24x7 on-call rotation is required and compensated.

Responsibilities

  • Build and scale ML-optimized HPC infrastructure, deploying and managing Kubernetes-based GPU/TPU superclusters across multiple clouds for high throughput and low-latency AI workloads.
  • Optimize infrastructure for AI/ML training in collaboration with cloud providers, focusing on cost efficiency, reliability, and performance using technologies like RDMA and NCCL.
  • Troubleshoot and resolve complex infrastructure bottlenecks, performance degradations, and system failures to ensure minimal disruption to AI/ML workflows.
  • Design and implement self-service tools and interfaces for researchers to monitor, debug, and optimize their training jobs.
  • Drive innovation in ML infrastructure by translating emerging researcher needs (e.g., JAX, PyTorch, distributed training) into robust, scalable solutions.
  • Champion best practices in observability, automation, and infrastructure-as-code (IaC) to ensure system maintainability and resilience.
  • Share expertise through code reviews, documentation, and cross-team collaboration to foster knowledge transfer and engineering excellence.

Requirements

  • Deep expertise in ML/HPC infrastructure, including GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and HPC environments.
  • Proven ability to deploy, manage, and troubleshoot cloud-native Kubernetes clusters at scale for AI workloads.
  • Proficiency in Python for ML tooling and Go for systems engineering.
  • Familiarity with Linux internals, RDMA networking, and performance optimization for ML workloads.
  • Experience working closely with AI researchers or ML engineers to solve infrastructure challenges.
  • Ability to identify bottlenecks, propose solutions, and drive impact in a fast-paced environment.

Benefits

  • Weekly lunch stipend of $75/£75 or equivalent.
  • Full health and dental benefits, including a separate budget for mental health.
  • RRSP matching, 401K, Pension Scheme.
  • 100% Parental Leave top-up for up to 6 months.
  • Annual enrichment benefits (Arts & culture, fitness/wellness, quality time, workspace improvement).
  • Education & learning stipend for conferences, courses, and coaching.
  • 6 weeks (30 working days) of paid vacation.
  • Budget for traveling to other offices if remote.
  • Annual company offsite.
  • $500 home office stipend.

About Cohere

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.