Software Engineer, GPU Infrastructure (HPC)
Remote • Canada • FullTime
Posted 6mo ago
Job Location
Canada
Tech Stack
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
Cohere is seeking a Staff Software Engineer to join our internal infrastructure team, responsible for building and operating world-class infrastructure and tools for training, evaluating, and serving Cohere's foundational AI models. You will work closely with AI researchers to support their AI workload needs on cutting-edge systems, focusing on stability, scalability, and observability. This role involves building and operating superclusters across multiple clouds, directly accelerating the development of industry-leading AI models. Participation in a 24x7 on-call rotation is required and compensated.
Responsibilities
- Build and scale ML-optimized HPC infrastructure, deploying and managing Kubernetes-based GPU/TPU superclusters across multiple clouds for high throughput and low-latency AI workloads.
- Optimize infrastructure for AI/ML training in collaboration with cloud providers, focusing on cost efficiency, reliability, and performance using technologies like RDMA and NCCL.
- Troubleshoot and resolve complex infrastructure bottlenecks, performance degradations, and system failures to ensure minimal disruption to AI/ML workflows.
- Design and implement self-service tools and interfaces for researchers to monitor, debug, and optimize their training jobs.
- Drive innovation in ML infrastructure by translating emerging researcher needs (e.g., JAX, PyTorch, distributed training) into robust, scalable solutions.
- Champion best practices in observability, automation, and infrastructure-as-code (IaC) to ensure system maintainability and resilience.
- Share expertise through code reviews, documentation, and cross-team collaboration to foster knowledge transfer and engineering excellence.
Requirements
- Deep expertise in ML/HPC infrastructure, including GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and HPC environments.
- Proven ability to deploy, manage, and troubleshoot cloud-native Kubernetes clusters at scale for AI workloads.
- Proficiency in Python for ML tooling and Go for systems engineering.
- Familiarity with Linux internals, RDMA networking, and performance optimization for ML workloads.
- Experience working closely with AI researchers or ML engineers to solve infrastructure challenges.
- Ability to identify bottlenecks, propose solutions, and drive impact in a fast-paced environment.
Benefits
- Weekly lunch stipend of $75/£75 or equivalent.
- Full health and dental benefits, including a separate budget for mental health.
- RRSP matching, 401K, Pension Scheme.
- 100% Parental Leave top-up for up to 6 months.
- Annual enrichment benefits (Arts & culture, fitness/wellness, quality time, workspace improvement).
- Education & learning stipend for conferences, courses, and coaching.
- 6 weeks (30 working days) of paid vacation.
- Budget for traveling to other offices if remote.
- Annual company offsite.
- $500 home office stipend.