ARC-AGI-3排行榜对不同AI系统在适应新型交互环境任务中的表现进行了评估对比 [1]。该排行榜的核心衡量指标为单任务成本与性能的关系,重点评估在有限计算资源约束下AI系统解决问题的效率 [1]。
排行榜设置了严格的成本限制,仅展示运行成本少于10,000美元的系统 [1]。参评系统包括多个类别:推理系统通过不同的推理级别展示思考时间增加如何影响性能,通常呈现渐近行为 [1];基础大模型包括GPT-4.5和Claude 3.7的单轮推理方案,这些模型不具备扩展推理能力 [1];此外还包括Kaggle竞赛提交的专门设计方案,在50美元计算预算限制下进行了120个评估任务 [1]。
根据排行榜的评分规则,无法产生完整测试输出的模型其余任务被标记为不正确 [1]。标记为"preview"的结果为非官方性质,可能基于不完整测试 [1]。
ARC-AGI从第1、2版本的被动流体智能测试演进至ARC-AGI-3的交互式环境适应挑战 [1],此次排行榜的发布为不同AI开发方案在资源效率与性能之间的权衡提供了量化对比基准 [1]。
The ARC-AGI-3 leaderboard provides a comparison of how different artificial intelligence systems perform on tasks involving adaptation to novel interactive environments, examining the relationship between computational cost efficiency and performance [1]. The leaderboard represents an evolution from earlier ARC-AGI versions, which assessed passive fluid intelligence, to a new interactive framework that tests how AI systems solve problems under limited computing resource constraints [1].
The primary metric used to evaluate systems is the relationship between per-task cost and performance, offering insight into resource efficiency [1]. Reasoning systems displayed on the leaderboard show multiple connection points representing different inference levels for the same model, demonstrating how increased thinking time affects performance, typically displaying asymptotic behavior [1]. Foundation models including GPT-4.5 and Claude 3.7 are represented through single-round inference without extended reasoning capabilities [1].
Kaggle competition submissions represent specially designed, efficient approaches and were evaluated across 120 tasks within a 50-dollar computational budget constraint [1]. The leaderboard includes only systems with total running costs below 10,000 dollars [1]. Models unable to generate complete test outputs have had their remaining tasks marked as incorrect [1]. Results labeled as "preview" are unofficial and may be based on incomplete testing [1].