Artificial Analysis发布了Intelligence Index v4.2版本更新,在其AI基准测试中加入AA-Briefcase和GDP.pdf两项新评估任务1。其中AA-Briefcase用于评估模型在多周知识工作项目上的表现,涉及数千个输入源文件1;GDP.pdf则是一项跨4,592页PDF的文档推理任务,包含100份PDF、十个领域、1,275条专家标准1。同时,该版本移除了GPQA Diamond评估1。
为了防止模型过拟合,Artificial Analysis将私有测试集的权重从v4.1版本的20%提升至40%,并计划在即将推出的Index v5中进一步提升这一比例1。在此次更新的评测中,Anthropic的Claude Fable 5.1在总排名中位列第一,OpenAI的GPT-6 Astra位列第二,相比前代GPT-5.6 Sol提升4分1。具体而言,GPT-6 Astra在GDP.pdf上得分33.2%,领先Claude Fable 5.1的26.2%;而Claude Fable 5.1则在AA-Briefcase上表现领先1。
Artificial Analysis has unveiled Intelligence Index v4.2, introducing significant updates to its AI benchmarking methodology 1. The new version incorporates two fresh assessment tasks—AA-Briefcase and GDP.pdf—while removing GPQA Diamond from the evaluation suite 1. To mitigate model overfitting, the company has increased the weighting of private test sets from 20% in v4.1 to 40%, with plans for further elevation in the upcoming Index v5 1.
In the latest rankings, Anthropic's Claude Fable 5.1 claims the top position, followed by OpenAI's GPT-6 Astra, which improved by four points compared to GPT-5.6 Sol 1. Performance differences emerge across specific tasks: GPT-6 Astra leads on GDP.pdf with a score of 33.2% against Claude Fable 5.1's 26.2%, while Claude Fable 5.1 maintains an advantage on the AA-Briefcase evaluation 1. AA-Briefcase assesses models on multi-week knowledge work projects involving thousands of source files, whereas GDP.pdf comprises a document reasoning challenge spanning 100 PDFs, ten domains, and 1,275 expert-validated benchmarks across 4,592 pages 1. The Index v4 release comes eight months after v4 debuted in January 2026, with Index v5 currently in development 1.
评论
还没有评论,欢迎留下第一条。