World Labs公司推出了名为Atlas的新一代世界模型,这是一个多模态自回归扩散变换器,能够原生处理文本、图像、视频和3D数据 1。该模型可执行包括相机控制生成、空间重建、时空模拟和图像生成等多项任务 1。
Atlas在多个关键任务上展现出卓越能力。该模型可生成最长1分钟、分辨率1440p的视频 1,并在3D重建任务中超越专业模型,仅需2-3张输入图像即可实现忠实重建 1。此外,Atlas支持像素级精确相机控制,用户可指定任意相机位置和角度 1。模型在空间重建和时空模拟中表现出色,适用于VFX、机器人和游戏等应用场景 1。该模型具有良好的可扩展性,性能随计算量增加而提升 1。
目前,Atlas已进入早期访问阶段,World Labs正在招募研究和工程人才以推进空间智能领域的发展 1。
World Labs has unveiled Atlas, a multimodal autoregressive diffusion transformer capable of natively processing text, images, video, and 3D data.1 The model performs multiple tasks including camera control generation, spatial reconstruction, spatiotemporal simulation, and image synthesis, outperforming specialized models on key benchmarks such as camera-conditional generation and 3D reconstruction.1
Atlas can generate video up to one minute in length at 1440p resolution.1 The system demonstrates particular strength in 3D reconstruction, achieving faithful reconstruction from just 2–3 input images while surpassing professional models in this domain.1 Users can exercise pixel-level precision in camera control, specifying arbitrary camera positions and angles within generated scenes.1 The model exhibits strong performance in spatial reconstruction and spatiotemporal simulation tasks, with applicability to visual effects, robotics, and gaming.1
The system has entered early access, with World Labs actively recruiting research and engineering talent to advance spatial intelligence capabilities.1 The model's performance scales with increased compute, indicating room for further improvement.1
评论
还没有评论,欢迎留下第一条。