OpenAI的GPT-6 Astra与Anthropic的Claude Fable 5/5.1在机器人臂操纵任务中的性能表现存在明显差异1。在「将红色积木放入碗中」的任务上,Astra展现了显著优势——在20次试验中成功19次,而Fable 5.1仅成功8次,Fable 5仅成功1次1。成本方面,Astra每次操作花费约0.94美元,不到Fable 5.1的2.12美元的一半1;在耗时上,Astra平均需要2.5分钟完成一次操作,而Fable 5.1则需要6.8分钟1。
然而在「将蓝色拼图片放入相应凹槽」的任务中,两个模型的表现相当,Astra和Fable 5.1均在20次试验中仅完成2次1。该任务上Astra的单次成本为1.36美元,略低于Fable 5.1的2.18美元1。
这项对比测试进行于2026年9月4日1,但存在若干限制条件:Astra的试验在Fable测试之后两天进行,碗任务使用了不同的机器人系统,且评分过程中操作者已知所测试模型的身份,存在潜在偏差1。
OpenAI's GPT-6 Astra outperformed Anthropic's Claude Fable 5.1 in a comparative test of robot arm manipulation abilities, though results varied significantly depending on task complexity 1. In a task requiring placing red blocks into a bowl, Astra achieved a 19 out of 20 success rate compared to Fable 5.1's 8 out of 20 and Fable 5's 1 out of 20 1. Astra also demonstrated substantial efficiency advantages in the same task, completing each attempt in an average of 2.5 minutes at a cost of $0.94, compared to Fable 5.1's 6.8 minutes and $2.12 per trial 1.
However, performance gaps narrowed considerably on a more complex task involving fitting blue puzzle pieces into corresponding slots 1. Both Astra and Fable 5.1 completed only 2 out of 20 attempts in this puzzle task, though Astra maintained a cost advantage at $1.36 per attempt versus Fable 5.1's $2.18 1. The testing was conducted on September 4, 2026, with notable methodological limitations: Astra trials took place two days after Fable testing, the bowl task employed different robotic systems for each model, and scorers knew which model was being tested, introducing potential evaluation bias 1.
评论
还没有评论,欢迎留下第一条。