觉
AI觉醒星球
Awakening is here
Knowledge File / 全球热点解读
2026-09-02 23 浏览 公开

World Labs 发布 Atlas:一个仅凭几张照片即可生成、重建并模拟 3D 世界的单一 AI 模型

World Labs 发布世界模型 Atlas,能从少量图像生成、重建和模拟 3D 场景。它训练于文本、图像、视频和 3D 数据,支持相机控制视频生成、少样本空间重建,并可作为机器人仿真工具,声称性能超越专用模型。

SOURCE / 全球热点解读 MIN / 9 ACCESS / 公开 POST / 2026-09-02 20:28:36

原贴

查看原文
作者:Jonathan Kemper 来源站点:the-decoder.com 原贴时间:

原文

World Labs, co-founded by AI researcher Fei-Fei Li, has announced Atlas, a world model that generates, reconstructs, and simulates 3D scenes from just a few images. The company claims it beats specialized models at their own tasks, which could make many of them unnecessary. Since its founding, World Labs has pursued the goal of "spatial intelligence" , the idea that AI should understand 3D space the way humans do. Atlas is the company's first model built to do that at scale. Rather than producing flat images or video clips, it grasps how a scene looks from any angle and how it changes over time. World Labs describes Atlas as an omni-model trained from scratch on text, images, video, and 3D data. Every input gets anchored to a specific position in 3D space rather than processed as a flat sequence. The company calls this shared spatial understanding "spatial context," and it's what the model uses to generate each new frame or viewpoint. According to World Labs, this anchoring separates Atlas from pure language or video models. Fei-Fei Li laid out this exact problem in a November 2025 essay. Current multimodal language models and video diffusion models break data into one- or two-dimensional sequences, she argued, which makes even simple spatial tasks needlessly hard. What's needed are architectures that organize tokenization, context, and memory in a 3D- or 4D-aware way. For camera-controlled generation, Atlas takes one or more images and produces new views at freely chosen camera positions and angles. Camera movement is passed as a direct geometric input rather than described through text prompts, as many video models require. The model outputs up to one minute of video at 1440p. Users can control every shot themselves instead of "pulling the lever on a slot machine," as World Labs put it, drawing a line between controlled generation and random output. For spatial reconstruction, Atlas rebuilds real scenes from as few as one to several dozen input images without special capture equipment. The more images it receives, the less it has to fill in from its own knowledge. With just two or three images, Atlas delivers faithful results and outperforms specialized 3D models, according to World Labs. It can also handle over a hundred inputs. In one demo, the model progressively assembles Stanford's Main Quad from two to 25 ground-level photos and generates aerial views far above the campus. This is where existing models tend to fall apart. In a comparison within the OpenWorldLib framework , systems like VGGT and InfiniteVGGT showed geometric inconsistencies and blurry textures as soon as the camera moved significantly. Atlas can output results as actual 3D data, not just images or video, because it processes depth information alongside RGB. Supported formats include point clouds and 3D Gaussian splats , which build a scene from many small spatial data points that can be viewed smoothly from any angle. This matches the representation used in Marble, the company's existing product . As a simulator, Atlas models space and time together. From footage captured by just a few cameras, it can produce a "bullet time" effect that freezes a scene and lets users view it from otherwise impossible angles. The demo footage was shot with a handful of smartphones and action cameras, not professional gear. For robotics, Atlas serves as a real-to-sim tool. It reconstructs a room and generates the image and depth data that a simulated robot's sensors would see along its path. From just a few photos, users can simulate and vary grasping and movement tasks by swapping out objects, positions, lighting, or backgrounds. The goal is to produce diverse training data for robots without capturing every situation in the real world.

中文翻译

由 AI 研究者李飞飞联合创立的 World Labs 宣布了 Atlas,一个仅凭几张图像就能生成、重建和模拟 3D 场景的世界模型。该公司声称,它在专用模型各自的任务上击败了它们,这可能使其中许多模型变得不必要。自成立以来,World Labs 一直追求“空间智能”的目标,即 AI 应该像人类一样理解 3D 空间。Atlas 是该公司首个为此大规模构建的模型。它不生成平面图像或视频片段,而是理解场景从任何角度看是什么样子,以及它如何随时间变化。

World Labs 将 Atlas 描述为一个从零开始、在文本、图像、视频和 3D 数据上训练的全能模型。每个输入都被锚定到 3D 空间中的特定位置,而不是作为平面序列处理。该公司称这种共享空间理解为“空间上下文”,模型用它来生成每个新帧或视角。据 World Labs 称,这种锚定将 Atlas 与纯语言或视频模型区分开来。

李飞飞在 2025 年 11 月的一篇文章中阐述了这个问题。她认为,当前的多模态语言模型和视频扩散模型将数据分解为一维或二维序列,这使得即使是简单的空间任务也变得不必要地困难。需要的是以 3D 或 4D 感知方式组织分词、上下文和记忆的架构。

对于相机控制的生成,Atlas 接收一张或多张图像,并在自由选择的相机位置和角度生成新视图。相机运动作为直接的几何输入传入,而不是像许多视频模型要求的那样通过文本提示描述。该模型输出长达一分钟的 1440p 视频。用户可以自己控制每个镜头,而不是像 World Labs 所说的那样“拉动老虎机的拉杆”,从而在受控生成和随机输出之间划清界线。

对于空间重建,Atlas 无需特殊捕捉设备,即可从少至一张、多至几十张输入图像中重建真实场景。它接收的图像越多,需要从自身知识中填补的就越少。据 World Labs 称,仅用两三张图像,Atlas 就能提供忠实的结果,并超越专用的 3D 模型。它还能处理超过一百个输入。在一个演示中,该模型从 2 到 25 张地面照片逐步组装出斯坦福大学的主庭院,并生成校园上方远处的鸟瞰视图。这正是现有模型容易崩溃的地方。在 OpenWorldLib 框架内的一项比较中,VGGT 和 InfiniteVGGT 等系统在相机大幅移动时立即显示出几何不一致和模糊纹理。

Atlas 可以将结果输出为实际的 3D 数据,而不仅仅是图像或视频,因为它将深度信息与 RGB 一起处理。支持的格式包括点云和 3D 高斯泼溅,它们从许多小型空间数据点构建场景,可以从任何角度平滑查看。这与该公司现有产品 Marble 中使用的表示相匹配。

作为模拟器,Atlas 将空间和时间一起建模。仅凭几台相机拍摄的素材,它就能产生“子弹时间”效果,冻结场景并让用户从其他方式不可能的角度查看。演示素材是用几部智能手机和运动相机拍摄的,而不是专业设备。

对于机器人技术,Atlas 充当真实到模拟的工具。它重建一个房间,并生成模拟机器人传感器沿其路径会看到的图像和深度数据。仅凭几张照片,用户就可以通过替换物体、位置、光照或背景来模拟和变化抓取与移动任务。目标是产生多样化的机器人训练数据,而无需在现实世界中捕捉每一种情况。

核心信息

World Labs 发布世界模型 Atlas,能从少量图像生成、重建和模拟 3D 场景。它训练于文本、图像、视频和 3D 数据,支持相机控制视频生成、少样本空间重建,并可作为机器人仿真工具,声称性能超越专用模型。

  • World Labs 发布世界模型 Atlas,能从少量图像生成、重建和模拟 3D 场景。它训练于文本、图像、视频和 3D 数据,支持相机控制视频生成、少样本空间重建,并可作为机器人仿真工具,声称性能超越专用模型。
  • 原贴提到:World Labs, co-founded by AI researcher Fei-Fei Li, has announced Atlas,
  • 来源:the-decoder.com

详细解读

这是什么信号

World Labs 发布 Atlas,不是又一版视频生成或 3D 重建工具,而是把“空间智能”作为统一建模目标。它试图用一个模型同时完成生成、重建和模拟,并且把相机位姿、深度、3D 数据都纳入同一套空间上下文。李飞飞在 2025 年 11 月的文章里已经点明问题:当前多模态模型把世界压成一维或二维 token 序列,导致简单空间任务也很困难。Atlas 是对这一判断的产品化回应。

为什么重要

如果 Atlas 的能力如宣传所言,它会冲击现有专用 3D 重建、视频生成和机器人仿真工具的分工。原文称其在少样本重建上超过专用 3D 模型,并支持点云、3D Gaussian Splats 等原生 3D 输出,这意味着下游可以继续编辑、渲染或用于仿真,而不是停留在像素视频。相机控制作为几何输入也把“可控生成”和“抽卡式生成”区分开。

对谁有价值

第一,3D 内容与游戏影视团队:少量照片就能重建场景并生成环绕视角,可降低资产制作和虚拟拍摄门槛。第二,机器人团队:real-to-sim 能快速构造带图像和深度的仿真环境,用于抓取和移动任务的数据增强。第三,空间智能研究者与创业者:Atlas 把 tokenization、context、memory 的 3D/4D 化作为架构问题,提供了可讨论的新范式。第四,AI 产品团队:如果模型开放,可围绕“空间上下文”做新交互。

可以怎么行动

关注 World Labs 是否开放 Atlas 的试用、API 或权重;如果可用,先用自有场景做小规模验证,对比现有 3D 重建管线在少样本、大角度变化下的几何一致性和纹理质量。机器人团队可以准备房间、物体、光照变化的数据集,测试 real-to-sim 生成的数据能否提升策略泛化。内容团队可以尝试用手机拍摄简单素材,验证相机控制生成和子弹时间效果的可用性。同时,把空间数据采集和 3D 资产整理纳入工作流,为这类模型做准备。

风险或限制

Atlas 的多数性能声明来自 World Labs 自己的对比和演示,尚未看到独立复现或大规模第三方评测。原文没有给出模型规模、推理成本、开放时间、许可方式和商业条款。少样本重建虽然声称忠实,但在复杂光照、反光、透明、动态物体和室外大范围场景中可能仍有局限。生成式重建还可能涉及隐私、版权和场景授权问题。机器人仿真到现实的差距也不会因为一个模型就消失,仿真数据仍需真实验证。对现有 3D 工具链而言,迁移成本和格式兼容性也是现实约束。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《World Labs 发布 Atlas:一个仅凭几张照片即可生成、重建并模拟 3D 世界的单一 AI 模型》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

World Labs 发布 Atlas:一个仅凭几张照片即可生成、重建并模拟 3D 世界的单一 AI 模型主要讲什么?

World Labs 发布世界模型 Atlas,能从少量图像生成、重建和模拟 3D 场景。它训练于文本、图像、视频和 3D 数据,支持相机控制视频生成、少样本空间重建,并可作为机器人仿真工具,声称性能超越专用模型。

这篇文章最值得关注的要点是什么?

World Labs 发布世界模型 Atlas,能从少量图像生成、重建和模拟 3D 场景。它训练于文本、图像、视频和 3D 数据,支持相机控制视频生成、少样本空间重建,并可作为机器人仿真工具,声称性能超越专用模型。;原贴提到:World Labs, co-founded by AI researcher Fei-Fei Li, has announced Atlas,;来源:the-decoder.com

这篇文章和哪些AI专题相关?

它适合放在AI工具、AI日报、Agent工作流专题里阅读。 关联原因:这篇内容命中「工具、模型」等主题信号。;这篇内容命中「热点解读」等主题信号。;这篇内容来自该专题长期覆盖的栏目。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI日报、每日AI日报、AI信号、热点解读、BuilderPulse这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 在 Confluent 上使用 IBM 时间序列模型实现实时智能 下一篇 ATV Big Air Tour 用 ChatGPT 把 3 天工作变成 3 小时