AI Engineer, Product
Remote • Paris • Full-time
Posted 7mo ago
Job Location
Paris
Tech Stack
Remote Work Policy
Fully remote
Employment Type
Full-time
Categories
Applied AI Engineer
About the job
Mistral AI is democratizing AI through high-performance, optimized, open-source models and products. We are a dynamic, collaborative team passionate about AI's potential to transform society, with teams distributed globally. We are seeking an AI Engineer to join a product team, focusing on improving AI-powered features through rigorous evaluation, prompt and orchestration design, and rapid experimentation. You will own your domain's AI quality end-to-end, defining, measuring, and shipping improvements in collaboration with our Science team.
Responsibilities
- Design and run evaluations for product areas, including reference tests, heuristics, and model-graded checks.
- Define and track key metrics such as task success, helpfulness, hallucination proxies, safety, latency, and cost.
- Own prompt and orchestration design, including writing, testing, and iterating on prompts and system prompts.
- Conduct A/B tests on prompts, models, and configurations, analyzing results to make rollout decisions.
- Set up observability for LLM calls, including structured logging, tracing, dashboards, and alerts.
- Manage model releases using canary and shadow traffic, sign-offs, and SLO-based rollback criteria.
- Improve core product area behaviors like memory policies, intent classification, routing, tool-call reliability, and retrieval quality.
- Create templates and documentation to enable other teams to author evaluations and ship safely.
- Partner with the Science team to diagnose regressions and lead post-mortems.
Requirements
- 3-4 years of experience, with backgrounds in ML engineering or software engineering with AI/ML production experience.
- Strong TypeScript or Python skills.
- Production LLM experience, including prompts, tool/function calling, and system prompts.
- Hands-on experience with evaluations and A/B testing, including metric design.
- Comfortable implementing directly in product code.
- Observability experience (logging, tracing, dashboards, alerting).
- Product mindset with the ability to form hypotheses, run experiments, interpret results, and ship features.
- Clear communication, autonomy, and a focus on production impact.
- Experience with safety systems (moderation, PII handling, guardrails) is ideal.
- Experience with release operations (canary/shadowing, automated rollbacks, experiment platforms) is ideal.
- Prior work on search ranking, chat systems, document AI, or audio ML features is ideal.