AMD推出了新一代Instinct MI455X GPU加速器,基于CDNA5架构[1]。相比前代MI355X系统支持8个GPU的规模,MI455X支持在单个Pod中扩展至72个GPU[1]。
单个MI455X包含256个Work Group Processors(WGPs)和8个Accelerator Complex Dies(XCDs)[1]。该加速器采用TSMC的CoWoS-L封装工艺[1],最大引擎时钟频率为2.4GHz[1]。在计算性能上,FP32矩阵/向量计算峰值达315 TFLOP,OCP MXFP4计算峰值达40.26 PFLOP[1]。
显存方面,单个MI455X配备12个HBM4堆栈,总计432GB显存[1],内存带宽达23.3TB/s[1]。
AMD同时推出了Helios机架级平台解决方案,用于大规模AI基础设施部署[1]。该平台采用基于UALink over Ethernet(UAEoL)的容错网络架构,支持单跳全互连通信[1]。Helios机架级平台可提供260TB/s的双向扩展带宽[1],当配置72个GPU时,可达到2.9 EF MXFP4性能或22.6 PF FP32性能[1]。
AMD has unveiled its latest-generation Instinct MI455X GPU accelerator, built on the CDNA5 architecture and designed to enable unprecedented scaling in AI infrastructure deployments. [1]
The MI455X represents a substantial leap in clustering capability, supporting up to 72 GPUs within a single Pod—a dramatic increase from the 8-GPU systems available in the predecessor MI355X lineup. [1] Each accelerator features 256 Work Group Processors and 8 Accelerator Complex Dies, paired with 432GB of HBM4 memory per GPU and memory bandwidth reaching 23.3TB/s. [1] The processors operate at a maximum engine clock frequency of 2.4GHz, delivering peak FP32 matrix and vector compute performance of 315 TFLOP per GPU, with OCP MXFP4 compute capability reaching 40.26 PFLOP. [1]
The MI455X employs TSMC's CoWoS-L packaging technology and utilizes a fault-tolerant network architecture based on UALink over Ethernet (UAEoL), enabling single-hop full-mesh communication across all connected accelerators. [1]
To support large-scale AI infrastructure deployment, AMD has introduced the Helios rackscale platform, which provides 260TB/s of bidirectional expansion bandwidth. [1] A fully populated 72-GPU Helios configuration delivers aggregate performance of 2.9 exaFLOPS in MXFP4 compute or 22.6 petaFLOPS in FP32 performance. [1]