Site Reliability Engineer, Inference Infrastructure
Remote • Toronto • FullTime
Posted 6mo ago
About the job
Cohere is seeking a Site Reliability Engineer to join the Model Serving team. This role is crucial for developing, deploying, and operating the AI platform that delivers Cohere's large language models via API endpoints. You will work closely with various teams to deploy optimized NLP models into production environments, ensuring low latency, high throughput, and high availability. The position also offers the opportunity to interact with customers and create customized deployments to meet their specific needs, contributing to the widespread adoption of AI.
Responsibilities
- Build self-service systems to automate service management, deployment, and operation, including custom Kubernetes operators for language model deployments.
- Automate environment observability and resilience to enable developers in troubleshooting and problem resolution.
- Ensure defined SLOs are met, including participation in an on-call rotation.
- Build strong relationships with internal developers and influence the Infrastructure team's roadmap based on feedback.
- Contribute to team development through knowledge sharing and an active review process.
Requirements
- 5+ years of engineering experience running production infrastructure at a large scale.
- Experience designing large, highly available distributed systems with Kubernetes and GPU workloads.
- Experience with Kubernetes development, production coding, and support.
- Experience with multi-cloud environments (GCP, Azure, AWS, OCI) and on-prem/hybrid serving.
- Experience in designing, deploying, supporting, and troubleshooting complex Linux-based computing environments.
- Experience in compute/storage/network resource and cost management.
- Excellent collaboration and troubleshooting skills for mission-critical systems.
- Grit and adaptability to solve complex, evolving technical challenges.
- Familiarity with computational characteristics of accelerators (GPUs, TPUs, custom) and their impact on inference latency and throughput.
- Strong understanding or working experience with distributed systems.
- Experience in Golang, C++, or other languages designed for high-performance scalable servers.
Benefits
- Weekly lunch stipend ($75/£75 or local equivalent).
- Full health and dental benefits, including a mental health budget.
- RRSP matching, 401K, Pension Scheme.
- 100% Parental Leave top-up for up to 6 months.
- Annual enrichment benefits (Arts & culture, fitness/wellness, quality time, workspace improvement).
- Education & learning stipend for conferences, courses, and coaching.
- 6 weeks of paid vacation (30 working days).
- Budget for traveling to other offices (for remote employees).
- Annual company offsite.
- $500 home office stipend.