Lead Member of Technical Staff, Inference Infrastructure

Remote San Francisco FullTime

Posted 3mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

Cohere is seeking a Lead Member of Technical Staff to join the Model Serving team. This role will provide technical leadership across multiple teams, driving the architecture and strategy for deploying optimized NLP models to production in low latency, high throughput, and high availability environments. You will be responsible for developing, deploying, and operating the AI platform that delivers Cohere's large language models through easy-to-use API endpoints. The ideal candidate will have a strong background in leading the design of high-performance, scalable, and reliable machine learning systems and will serve as a key point of contact for customers, leading the design of customized deployments and mentoring engineers.

Responsibilities

  • Provide technical leadership across multiple teams.
  • Drive the architecture and strategy for deploying optimized NLP models to production.
  • Develop, deploy, and operate the AI platform for delivering large language models.
  • Serve as a key point of contact for customers, leading the design of customized deployments.
  • Mentor engineers to raise the technical bar across the team.
  • Troubleshoot complex Linux-based computing environments at scale.
  • Own and optimize compute/storage/network resources and costs at an organizational level.

Requirements

  • 8+ years of engineering experience running production infrastructure at large scale.
  • Track record of technical leadership.
  • Experience leading the architecture and design of large, highly available distributed systems with Kubernetes and GPU workloads.
  • Deep expertise with Kubernetes development, production coding, and support.
  • Extensive experience across GCP, Azure, AWS, OCI, and multi-cloud on-prem/hybrid serving environments.
  • Ability to guide strategic infrastructure decisions.
  • Proven ability to lead the design, deployment, support, and troubleshooting of complex Linux-based computing environments at scale.
  • Experience owning compute/storage/network resource and cost management, including optimization strategies.
  • Exceptional collaboration and communication skills.
  • Experience mentoring engineers and leading cross-functional initiatives.
  • Grit and adaptability to solve and guide others through complex technical challenges.
  • Strong expertise in the computational characteristics of accelerators (GPUs, TPUs, custom accelerators) and leveraging them for latency and throughput improvements.
  • Deep knowledge of distributed systems, with experience establishing patterns and practices across engineering teams.
  • Proficiency in Golang, C++ or other high-performance scalable server languages.
  • Ability to set coding standards and conduct senior-level technical reviews.

Benefits

  • Weekly lunch stipend of $75/£75 or equivalent.
  • Full health and dental benefits.
  • Separate budget for mental health.
  • RRSP matching, 401K, Pension Scheme.
  • 100% Parental Leave top-up for up to 6 months.
  • Annual enrichment benefits (Arts & culture, fitness/wellness, quality time, workspace improvement).
  • Education & learning stipend for conferences, courses, and coaching.
  • 6 weeks of paid vacation (30 working days).
  • Budget for traveling to other offices.
  • Annual company offsite.
  • Co-working benefit for those not near an office.
  • Home office stipend of $500.

About Cohere

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.