一份针对软件工程任务的AI模型性能评测排行榜已发布,涵盖13个模型和4个Agent在Go、Java、Python、Rust、TypeScript等编程语言上的表现1。本次基准测试基于111个问题和来自65个开源仓库的数据进行评估1。
排行榜已进行多轮更新。2026年7月1日新增的模型包括GLM 5.2、DeepSeek-V4 Pro、DeepSeek-V4 Flash、MiMo V2.5 Pro、Qwen 3.6系列和Gemma 4 31B1。此前在2026年5月27日曾新增gpt-5.5系列、Claude Opus 4.7和Kimi K2.61。评测框架在2025年9月4日引入了Cost per Problem和Tokens per Problem两项新指标1。所有排行榜问题的Docker镜像和HuggingFace数据集已于2025年7月11日发布1。
A new performance evaluation leaderboard has been published assessing artificial intelligence models on software engineering (SWE) tasks, encompassing 13 models and 4 agents across multiple programming languages 1. The benchmark evaluates performance on 111 problems drawn from 65 open-source repositories and covers Go, Java, Python, Rust, and TypeScript 1.
The leaderboard has undergone continuous updates with the addition of multiple new models over recent periods 1. As of July 1, 2026, the ranking introduced GLM 5.2, DeepSeek-V4 Pro, DeepSeek-V4 Flash, MiMo V2.5 Pro, Qwen3.6 series, and Gemma 4 31B 1. Prior to that, on May 27, 2026, the leaderboard added gpt-5.5 series, Claude Opus 4.7, and Kimi K2.6 1. The benchmark has also retired certain older model versions from the rankings 1.
In addition to model performance metrics, the leaderboard introduced new evaluation dimensions on September 4, 2025, including Cost per Problem and Tokens per Problem 1. Supporting the transparency and reproducibility of the evaluation, the project released Docker images for all benchmark problems and a corresponding HuggingFace dataset on July 11, 2025 1.
评论
还没有评论,欢迎留下第一条。