Model Evaluation Jobs
2 open roles mentioning Model Evaluation
Forward Deployed Product Manager, Enterprise
Scale AI
Scale is seeking a Forward Deployed Product Manager (FDPM) to drive the success of enterprise AI deployments. This role is embedded with customers, focusing on achieving real production outcomes and translating operational realities into actionable product insights. Unlike a traditional roadmap PM or solutions engineer, the FDPM owns product outcomes within a portfolio of enterprise accounts, building trust with senior stakeholders, guiding deployments to production, and identifying product versus execution bottlenecks. The ideal candidate can differentiate between stated customer needs, underlying problems, and optimal platform solutions, having previously shipped products into large organizations.
$206k - $300k
Model Behavior Architect- Function Calling
Mistral AI
As a Model Behavior Architect on the Function Calling team, you will be at the forefront of defining and measuring how Large Language Models (LLMs) utilize tools, invoke functions, and orchestrate complex agentic workflows. We are seeking individuals with a strong background in engineering, machine learning, and LLMs, who possess expertise in model evaluation, policy writing, and creating evaluation pipelines specifically for tool use and function calling. Your primary responsibility will be to collaborate closely with our Science team to establish benchmarks for function calling performance, covering aspects such as accurate parameter selection, schema adherence, multi-step tool orchestration, error recovery, and agentic reasoning. This role is ideal for those passionate about tackling cutting-edge, open-ended research challenges and translating insights into best-in-class models.