Google DeepMind·· 2025-12-09精选AI 评分78
Google DeepMind 发布 FACTS Benchmark Suite 评测大语言模型事实性
FACTS Benchmark Suite: Systematically evaluating the factuality of large language models
AI 导读
Google DeepMind 联合 Kaggle 发布 FACTS Benchmark Suite,新增 Parametric、Search、Multimodal 三项基准,并更新 FACTS Grounding Benchmark v2。
推荐理由
FACTS Benchmark Suite 将内部知识、网页搜索、图像问答和上下文 grounding 分开评测,并公开 3,513 个样例,便于比较模型在不同事实性任务上的表现。
来源:Google DeepMind · deepmind.google