近期大语言模型发布后,社交媒体上出现大量相同类型的演示,包括Minecraft重建、MS Paint绘画和SVG动画等1。然而评论指出,这些演示基准测试存在根本缺陷:它们是固定、众所周知的目标,可被AI实验室通过针对性优化完美解决,因此衡量的是准备程度而非真实能力1。研究表明,实验室可在8周内对这些测试进行过度拟合1。
验证这一问题的证据已浮现:Thinking Machines的Inkling Small模型参数不足旗舰版本的三分之一,却在Artificial Analysis上得分相近,在Humanity's Last Exam、GPQA Diamond和SciCode上表现甚至更优1。这表明参数规模较小的模型仍能通过优化达到相似或更好的演示效果,进一步证明演示基准难以准确反映模型的实际能力。
为应对这一问题,部分评估方案已采取措施:LiveBench采用轮转问题机制,ARC-AGI保持测试集私密,Humanity's Last Exam则保留部分测试的隐藏性1。评论建议,演示基准适合用于营销和社交内容,但不应被用于真实能力评分1。
Recent demonstrations of large language models like GPT Astra have sparked critical discussion about the validity of using fixed, publicly known tasks to evaluate AI capabilities.1 Critics argue that these demonstration benchmarks—which typically include challenges such as Minecraft reconstruction, MS Paint drawing, and SVG animation—represent a fundamental flaw in how model performance is being measured.1 Because these targets are static and well-known, AI laboratories can optimize their systems specifically to solve them, meaning the benchmarks measure preparation rather than genuine ability.1
The problem becomes evident when examining model comparisons. Thinking Machines' Inkling Small model, which contains less than one-third the parameters of flagship versions, achieves comparable scores on Artificial Analysis while outperforming larger models on Humanity's Last Exam, GPQA Diamond, and SciCode.1 This discrepancy suggests that demonstration benchmarks may not accurately reflect real-world capabilities.1 To address these concerns, some organizations have adopted alternative evaluation approaches: LiveBench rotates its questions, ARC-AGI maintains a private test set, and Humanity's Last Exam reserves portions of its tests from public view.1 Observers contend that while demonstration benchmarks serve valuable purposes in marketing and social media engagement, they should not be relied upon for assessing genuine model performance.1
评论
还没有评论,欢迎留下第一条。