GitHub推出Project HydraFusion研究预览版,这是一套通过多模型运行时编排实现代码生成的AI系统1。该系统能根据任务从多个提供商的模型中动态选择执行工作流,支持单模型、级联升级和评审修订三种执行模式1。
在多项基准测试中,HydraFusion展现出显著的成本优势和竞争力的质量表现。在TerminalBench 2.1基准测试中,该系统相比Claude Opus 5提升4.9个百分点的验证任务质量,同时降低67%的估计成本1。在DeepSWE基准测试中,HydraFusion与Opus 5的质量差距为1.5个百分点,成本降低36%1。在CheckpointBench基准测试中,两者质量差异仅为0.1个百分点,而成本下降65%1。
HydraFusion现已作为研究预览版集成进GitHub Copilot,供开发者使用1。
GitHub has unveiled Project HydraFusion, a research preview system that orchestrates multiple AI models to generate code while maintaining quality and reducing costs 1. The platform dynamically selects execution workflows from models across different providers—operating in single-model, cascade, or critique modes depending on task requirements 1.
In offline benchmark testing, HydraFusion demonstrates competitive performance against Claude Opus 5 with significant cost savings 1. On TerminalBench 2.1, the system outperforms Opus 5 by 4.9 percentage points on verification tasks while cutting estimated costs by 67% 1. The DeepSWE benchmark shows HydraFusion trailing Opus 5 by 1.5 percentage points but delivering 36% cost reduction 1, while CheckpointBench results show a negligible 0.1 percentage point quality difference alongside 65% cost savings 1.
The technology has been integrated into GitHub Copilot, making it available to developers as a research preview 1.
评论
还没有评论,欢迎留下第一条。