英国人工智能安全研究所(UK AISI)与美国AI标准创新中心(CAISI)近日联合发布了对月之暗面公司Kimi K3模型网络能力的初步评估报告。[1]
评估结果
在漏洞利用基准测试(ExploitBench)中,Kimi K3获得32%的得分,超过了同类开源模型GLM-5.2的24%。[1]该测试涵盖了41个2023年后发现的V8引擎漏洞。[1]
但在高级网络攻击模拟测试"The Last Ones"(TLO)中,Kimi K3的表现明显低于顶尖美国模型。[1]根据测试数据,Kimi K3在该攻击链中平均完成至第17步(总共32步),而美国顶尖模型平均完成28.5步。[1]在10次尝试中,Kimi K3仅成功完成1次。[1]
在任意代码执行(ACE)能力方面,Kimi K3在41个样本测试中达成0次,而最顶尖模型平均达成20次。[1]评估结果表明Kimi K3未能在任何任务中实现任意代码执行。[1]
产品发布计划
Kimi K3已于2026年7月16日发布,计划于2026年7月27日发布开源版本。[1]
The UK Artificial Intelligence Safety Institute (UK AISI) and the US Center for AI Safety and Innovation (CAISI) have jointly evaluated the cyber capabilities of Kimi K3, a model developed by Moonshot AI. [1]
The assessment measured performance across multiple cybersecurity benchmarks. On the ExploitBench vulnerability exploitation benchmark, Kimi K3 achieved a score of 32%, outperforming the open-source GLM-5.2 model, which scored 24%. [1] However, in advanced attack simulation testing known as "The Last Ones" (TLO), Kimi K3 performed below leading American models. [1]
Notably, Kimi K3 failed to achieve arbitrary code execution (ACE) in any test, registering zero successful instances across 41 samples, while leading American models averaged 20 successful ACE instances. [1] In the TLO attack chain assessment, Kimi K3 completed an average of step 17 out of 32 total steps, compared to an average of step 28.5 for top American models. [1] The model achieved only one successful completion in 10 attempted runs of the TLO scenario. [1]
The ExploitBench testing covered 41 V8 engine vulnerabilities discovered after 2023. [1] Kimi K3 was released on July 16, 2026, with an open-source version scheduled for release on July 27, 2026. [1]