Flux公司推出了新一代多模态基础模型FLUX 3,支持图像、视频和音频的联合学习与生成[1]。该模型能够生成长达20秒的视频和音频内容[1],并可合成和编辑各种风格的图像[1],同时具备动作预测能力[1]。
在初步性能评估中,FLUX 3相比多个竞品表现突出。与Runway Gen-4.5对比中被优选77%[1],与Luma Ray 3.2对比中被优选93%[1],与Grok Imagine Video对比中被优选69%[1],与Kling v3 Pro对比中被优选60%[1]。
目前FLUX 3 Video已进入早期访问阶段[1],FLUX 3 Image将在随后数周开放早期访问[1]。Flux计划提供多种访问方式,包括API、私有权重访问和开源权重访问[1]。此外,该公司与mimic robotics合作开发了FLUX-mimic模型用于机器人学习应用[1]。
Flux has introduced Flux 3, a new multimodal foundation model capable of jointly learning across images, video, and audio [1]. The model is now available in early access, with video generation functionality already accessible to users [1].
The system can generate videos and audio lasting up to 20 seconds, synthesize and edit images across various styles, and perform action prediction tasks [1]. In preliminary comparative evaluations, Flux 3 demonstrated strong performance against competing systems: it was preferred over Runway Gen-4.5 in 77% of comparisons, over Luma Ray 3.2 in 93% of comparisons, over Grok Imagine Video in 69% of comparisons, and over Kling v3 Pro in 60% of comparisons [1].
Flux 3 Video is currently available during the early access phase, while Flux 3 Image will become available in the coming weeks [1]. The company has announced plans to provide multiple access methods, including API access, private weight access, and open-source weight availability [1].
Additionally, Flux is collaborating with Mimic Robotics to develop Flux-Mimic, intended for machine learning applications in robotics [1].