evaluation Jobs
2 open roles mentioning evaluation
Product Manager, Gen AI
Scale AI
Scale AI is seeking Product Managers to join its GenAI organization, focusing on building the data infrastructure that powers advanced AI. These roles involve shaping systems, tooling, and experiences for a two-sided marketplace connecting AI labs and enterprises with a global network of contributors. You will work on high-impact, technically complex problems at the frontier of AI, owning product areas end-to-end from strategy to execution. The positions require deep cross-functional collaboration with engineering, design, data science, operations, and other stakeholders in a fast-paced, growth-stage environment.
Model Behavior Architect- Safety
Mistral AI
As a Model Behavior Architect, you will be at the forefront of defining and measuring LLM behavior. We are seeking individuals with a background in engineering, machine learning, and large language models, who possess expertise in model evaluation, policy writing, and creating evaluation pipelines for complex tasks. Your role will involve close collaboration with our Science team to establish standards for Reasoning, Audio, Alignment, Tools, and other frontier initiatives. This is an opportunity to engage with cutting-edge, open-ended research challenges and translate your insights into superior models.