AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-07-27 0 浏览 会员

脑电波会是物理AI的下一个解锁点吗?

物理AI发展正面临真实世界训练数据稀缺的瓶颈,Encord等初创公司尝试通过脑电波等新模态制造数据,探索突破之路。

SOURCE / AI小生意项目库 MIN / 4 ACCESS / 会员 POST / 2026-07-27 08:19:14

原贴

查看原文
作者:Tim Fernholz 来源站点:techcrunch.com 原贴时间:

原文

The frontier of physical AI is a Jenga game in a warehouse in San Leandro, California. That warehouse is occupied by Encord , a company that builds data tooling used to train AI models. Andrew Ceja is a pilot—the company’s term for its robotic trainers—and he’s carefully pulling wooden blocks from a tottering tower while wearing a headset with a camera that tracks what he sees. That alone is fairly common for collecting robot training data, but this headset includes sensors that measure his brain waves as he carefully disassembles the block tower. Encord is one of a small but growing number of startups betting that the next real constraint on humanoid and warehouse robotics won’t be model architecture but instead the sheer scarcity of real-world physical training data. Rather than just helping robotics companies manage the data they have, Encord is building a business around manufacturing the data they don’t. The brain wave headset Ceja is wearing was built by Zander Labs , a German neuroscience startup that’s betting measuring brain activity — to deduce mental states like error, intent and surprise — can create a more useful data set to train models. Encord’s work with Zander is currently a trial run; Encord says the goal is to build an initial brain wave-tagged data set, run it through customer robotics models, and evaluate whether it actually improves performance before deciding whether to scale it up. Lucas Gehrke, a Zander neuroscientist supervising the work, says that the amount of brain activity used at any point during a given task offers clues for model builders trying to figure out when they need to deploy their highest-effort models. This is the “bleeding edge” of the effort to solve the robotics data bottleneck, according to Vineeth Velmurugan, Encord’s head of robot learning. A veteran of OpenAI’s robot lab and Berkshire Grey, the warehouse automation firm, Velmurugan joined Encord to build the company’s internal data-creation team. Encord was founded to help companies building machine-vision applications annotate data and evaluate models. As their customers—Velmurugan says they work with many leading robotics firms but that he’s not authorized to name them—began to apply end-to-end learning to robotic manipulation tasks, executives realized they would have to produce training data themselves, rather than simply manage it. “The data simply does not exist,” Velmurugan said. The bet that generative AI can do for robots what it’s done for chatbots keeps running into this same wall. LLMs were built on the text of the entire internet, and more. Finding the same raw materials to teach neural networks about physical manipulation is challenging: self-driving car companies collect it themselves, but that’s hard to scale. Training from video can work, but it lacks the fidelity of real world data. Velmurugan says it will take a data set something like five times the size of YouTube’s video corpus to break through—a scale that helps explain why data-generation itself has become a business and not just a research problem. Companies building robot brains are now turning to two main sources: “Egocentric” video collected by workers wearing cameras, often augmented with additional camera angles and other metrics, and collecting data from robots operated remotely. Encord does both, drawing egocentric data from several factories around the globe, and using its San Leandro facility to experiment with new modalities, like brain waves, or collect data sets around specific skills for fine-tuning. When TechCrunch visited, pilots were using leader-follower rigs — paired robotic arms, one controlled directly by a human operator and one that mimics its movements —to create data about tasks like pouring coffee from a pot into mugs (very sloshy) and stacking poker chips. “Every humanoid company has asked us for these pieces,” Velmurugan says. Storage racks held cartons of fake flowers in vases, books, plastic vegetables, kitty litte

中文翻译

物理AI的前沿是加州圣莱安德罗仓库里的一场叠叠乐游戏。那个仓库由Encord公司占用,该公司构建用于训练AI模型的数据工具。Andrew Ceja是一名飞行员——公司对机器人训练师的称呼——他戴着带摄像头的耳机,仔细地从摇摇欲坠的塔上拉出木块,摄像头追踪他所看到的东西。仅这一点对于收集机器人训练数据来说相当常见,但这个耳机还包括传感器,在他小心拆解积木塔时测量他的脑电波。

Encord是少数但不断增长的初创公司之一,它们押注人形机器人和仓库机器人的下一个真正限制将不是模型架构,而是现实世界物理训练数据的极度稀缺。Encord不仅仅帮助机器人公司管理他们已有的数据,还在围绕制造他们没有的数据来构建业务。

Ceja佩戴的脑电波耳机由Zander Labs制造,这是一家德国神经科学初创公司,押注测量大脑活动——以推断错误、意图和惊讶等心理状态——可以创建更有用的数据集来训练模型。Encord与Zander的合作目前是试验性的;Encord表示,目标是构建一个初步的带有脑电波标记的数据集,通过客户的机器人模型运行,并评估它是否真的提高了性能,然后再决定是否扩大规模。

Zander的神经科学家Lucas Gehrke监督这项工作,他说,在给定任务中任何时刻使用的脑活动量,为试图弄清楚何时需要部署最高努力模型的模型构建者提供了线索。

Encord的机器人学习主管Vineeth Velmurugan表示,这是解决机器人数据瓶颈努力的“前沿”。Velmurugan是OpenAI机器人实验室和仓库自动化公司Berkshire Grey的资深人士,他加入Encord是为了建立公司的内部数据创建团队。

Encord成立的初衷是帮助构建机器视觉应用的公司注释数据和评估模型。当他们的客户——Velmurugan说他们与许多领先的机器人公司合作,但他未被授权透露名字——开始将端到端学习应用于机器人操作任务时,高管们意识到他们必须自己生成训练数据,而不是简单地管理数据。

“这些数据根本不存在,”Velmurugan说。生成式AI能为机器人做到像为聊天机器人所做的事情,这一赌注不断撞上同一堵墙。LLM是建立在整个互联网文本之上的,甚至更多。寻找同样的原材料来教神经网络物理操作是具有挑战性的:自动驾驶汽车公司自己收集数据,但这很难扩展。从视频训练可能有效,但它缺乏现实世界数据的保真度。Velmurugan表示,要取得突破,需要一个大约是YouTube视频语料库五倍大小的数据集——这个规模有助于解释为什么数据生成本身已经成为一门生意,而不仅仅是研究问题。

构建机器人大脑的公司现在转向两个主要来源:由佩戴摄像头的工人收集的“第一人称”视频,通常辅以额外的摄像机角度和其他指标,以及从远程操作的机器人收集数据。Encord两者都做,从全球多个工厂获取第一人称数据,并利用其圣莱安德罗设施试验新模态,如脑电波,或围绕特定技能收集数据集以进行微调。

TechCrunch参观时,飞行员们正在使用主从装置——配对的机械臂,一个由人类操作员直接控制,另一个模仿其动作——来创建数据,涉及从壶里倒咖啡到杯子(非常晃荡)和堆叠扑克筹码等任务。“每个人形机器人公司都向我们要过这些,”Velmurugan说。存储架上放着花瓶里的假花、书、塑料蔬菜、猫砂(原文截断)

核心信息

物理AI发展正面临真实世界训练数据稀缺的瓶颈,Encord等初创公司尝试通过脑电波等新模态制造数据,探索突破之路。

  • 物理AI发展正面临真实世界训练数据稀缺的瓶颈,Encord等初创公司尝试通过脑电波等新模态制造数据,探索突破之路。
  • 原贴提到:The frontier of physical AI is a Jenga game in a warehouse in San Leandr
  • 来源:techcrunch.com
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 组织:已统治地球的超级智能体 下一篇 转售和欺诈驱动的地下LLM令牌转售市场内幕