研究人员发表论文探索利用MUD(文本冒险游戏)来评估大型语言模型,总耗资仅99美元[1]。研究对四个行为维度进行评分,其中两个依赖LLM分类器[1]。
实验发现了LLM评判器的严重可靠性问题。两个评判器之间的模型一致性范围从85%至22%[1],在探针检测上的aggregate kappa值仅为0.04[1]。这个结论本身比LLM排名结果更具价值[1]。当移除两个LLM分类器维度后,一个前沿模型的排名下降了六位[1]。
研究中每个模型仅运行50次[1]。研究人员强调这项工作仅构成概念验证而非正式基准[1],已将全部数据和代码开源,论文和数据采用CC BY 4.0许可,代码采用MIT许可[1]。相关研究成果已上传至Zenodo[1]。
Researchers have published a paper exploring the use of MUDs—text-based adventure games—as a novel evaluation method for large language models, completing the project in months with costs totaling merely $99 in API credits [1]. The study represents a proof of concept rather than a formal benchmark [1].
The evaluation framework assessed models across four behavioral dimensions, two of which relied on LLM-based classifiers [1]. When these two LLM classifier dimensions were removed from the analysis, one leading model's ranking dropped by six positions [1]. This finding underscores a critical insight: the reliability issues inherent in using LLM evaluators proved more significant than the ranking variations themselves [1].
The research revealed substantial inconsistencies between evaluators. Agreement between two classifiers on model rankings ranged from 85 percent down to just 22 percent [1], while the aggregate kappa value—a measure of inter-rater reliability—reached only 0.04 on detection probes [1]. Each model underwent evaluation in only 50 runs [1].
The researchers have made their work openly accessible, releasing all data and code under the CC BY 4.0 license for the paper and data, and the MIT license for the code [1]. The full materials are available through Zenodo at https://doi.org/10.5281/zenodo.21386663 [1].