Applied AI Engineer, Site Reliability Engineer - EMEA
Remote • Paris • Full-time
Posted 2mo ago
Job Location
Paris
Tech Stack
Remote Work Policy
Fully remote
Employment Type
Full-time
Categories
Applied AI Engineer
About the job
Mistral AI is seeking a founding engineer for its Applied AI Site Reliability Engineering (SRE) sub-team. This role is crucial for building and operating a framework that ensures the reliability and sustainability of Mistral's AI solutions across all customer accounts, whether hosted by Mistral or the customer. You will operate in four key modes: BUILD (designing for a fleet of platforms, proactive reliability, authoring runbooks, implementing observability), RUN (operating Tier-1 customer environments, ensuring SLO compliance, managing incidents), ENABLE (productizing deployment, security, and scaling of Applied AI solutions), and SECURE (owning security operations, leading CVE response, and implementing supply-chain integrity controls). This is a framework-first, fleet management role focused on structurally solving problems for all customers, not just individual ones. The team values people and outputs, direct feedback, low ego, and high standards in a fast-paced, unstructured environment.
Responsibilities
- Design and build frameworks for reliable and sustainable delivery of Mistral's AI solutions.
- Operate and maintain Tier-1 customer environments, ensuring Service Level Objective (SLO) compliance.
- Manage on-call duties and lead incident response efforts, including conducting blameless post-mortems.
- Develop and productize processes for deploying, securing, and scaling Applied AI solutions.
- Own the security operations layer for customer-side deployments, including CVE response and supply-chain integrity.
- Implement observability stacks for monitoring and debugging production systems.
- Automate provisioning, security guardrails, and configuration baselines.
- Ensure robust security practices, including secure-SDLC, CVE response, and supply-chain integrity controls.
Requirements
- 5+ years of experience in SRE, Production Engineering, or DevOps with a track record of shipping tooling.
- Strong fluency in multi-tenant Kubernetes, including namespace segmentation, network policy, RBAC, and admission control.
- Experience with incident response, blameless post-mortems, and a runbook-first mindset.
- Proficiency with observability stacks such as Prometheus, Grafana, OpenTelemetry, Loki, Tempo, or Signoz.
- Experience with Infrastructure as Code tools like Terraform or Ansible.
- Proficiency in Python and/or Golang for tooling and automation.
- A strong security mindset, treating secure-SDLC, CVE response, and supply-chain integrity as core reliability properties.
- Excellent written communication skills for runbooks, post-mortems, and incident communications.
- Ability to operate with high autonomy in an ambiguous, fast-paced environment.
- Solid understanding of Linux internals, networking debugging, and distributed systems fundamentals.