Black Forest Labs 与 mimic robotics 宣布推出 FLUX-mimic,这是一款基于 FLUX 3 多模态基础模型的视频-动作预测系统 [1]。FLUX 3 在统一训练框架中可同时生成图像、视频和音频,并能预测机器人动作 [1]。
在模型开发中,视频预测在联合训练的总计算成本中占比超过 95% [1]。研究团队通过 Self-Flow 实验发现,动作预测相比无此模块的视频模型能节省约一半的训练步数 [1]。添加动作预测后,模型在 3500 步内恢复了视频生成质量,性能下降随后完全消失 [1]。
在推理性能上,FLUX-mimic 的 backbone 在单个 NVIDIA RTX 5090 GPU 上可在 80 毫秒内完成从输入到世界表示的处理 [1]。mimic 优化的完整部署系统实现了 101 毫秒的机器人反应时间 [1]。相比视觉-语言-动作模型,mimic-video 论文报告的样本效率提升达到 10 倍 [1]。
该系统已在奥迪生产线上开展测试部署,用于处理零件装配、电子控制单元插装、部件组装和柔性材料处理等任务 [1]。奥迪的 Christoph Schneider 表示:"这些机器人解决了传统机器人不可能处理的复杂软体操作任务" [1]。
Black Forest Labs and mimic robotics have unveiled FLUX-mimic, a next-generation video-action model built on the FLUX 3 multimodal foundation model [1]. FLUX 3, trained jointly, generates images, video, and audio while simultaneously predicting robot actions [1].
The system has been deployed for testing on Audi production lines, where it handles complex tasks traditionally difficult for conventional robotics, including soft material manipulation and precision assembly operations [1]. According to Christoph Schneider of Audi, "these robots solve complex soft-body manipulation tasks that traditional robots cannot possibly handle" [1]. Deployment tasks encompass parts assembly, electronic control unit insertion, component assembly, and flexible material handling [1].
Performance and Efficiency
During joint training, video prediction accounts for over 95% of total computational cost [1]. When action prediction was added, the model recovered full video generation quality within 3,500 training steps, with performance degradation subsequently disappearing [1]. The FLUX-mimic backbone processes input to world representation in 80 milliseconds on a single NVIDIA RTX 5090 GPU [1], while mimic's optimized full deployment system achieves a robot reaction time of 101 milliseconds [1].
Training efficiency gains are notable: the Self-Flow approach reduced training steps by half compared to video models without Self-Flow [1], and the mimic-video approach demonstrated a tenfold improvement in sample efficiency relative to vision-language-action models [1].