研究人员通过系统实验检验了一个假设:顶级AI实验室是否针对"生成pelican骑自行车"这一著名基准任务对模型进行了优化。[1]
该实验测试了7个前沿AI模型——GPT-5.6 Terra、Claude Sonnet 5、Gemini 3.5 Flash、Grok 4.5、Qwen3.7-Max、GLM-5.2和DeepSeek V4 Pro。[1]研究人员让这些模型各生成了1,008张SVG图像,涵盖8种动物与6种交通工具的组合,每种组合包含3个样本。[1]
分析结果表明,pelican和bicycle都不是各模型表现最优的对象。在8种动物的绘制质量排名中,pelican位列第6;在6种交通工具中,bicycle排名倒数第二。[1]pelican与bicycle的组合在全部48种组合中则排到第42位。[1]
统计检验未能找到针对性优化的证据。固定效应回归分析显示,pelican效应的p值最小为0.25;在bicycle效应中,仅有Gemini达到p<0.05的显著性水平,但这一结果无法通过多重比较修正。[1]
即使在看似异常的细节上,也未观察到有针对性的调整。所有21张pelican-bicycle组合的图像均朝向右边,但在整体数据集中,60%的图像同样朝向右边。[1]
这项实验的成本相对低廉,仅需约80美元的API费用。[1]
Researcher Dylan Castillo conducted an experiment to investigate whether leading artificial intelligence laboratories have optimized their models to excel at generating images of pelicans riding bicycles, a phenomenon he termed "pelicanmaxxing." [1]
The study tested seven frontier AI models—GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro—by generating a total of 1,008 SVG images across eight animal types and six transportation methods, with three samples per combination. [1]
The results provided little support for the pelicanmaxxing hypothesis. Pelican ranked sixth among the eight animals tested, while bicycle ranked fifth out of six transportation methods. [1] Across all 48 possible animal-vehicle combinations, the pelican-bicycle pairing ranked 42nd. [1]
Statistical analysis using fixed-effects regression revealed no meaningful evidence of targeted optimization. All pelican effect p-values were at minimum 0.25, and among bicycle effects, only Gemini achieved p<0.05—a result that could not withstand multiple comparison correction. [1]
One observation stood out: all 21 pelican-bicycle images were oriented rightward, yet 60 percent of images across the entire dataset faced the same direction, suggesting no anomalous bias specific to the combination. [1]
The experiment was conducted on a modest budget of approximately 80 dollars in API costs. [1]