觉
AI觉醒星球
Awakening is here
Knowledge File / 全球热点解读
2026-08-26 5 浏览 公开

AgentHands:在XR中为空间锚定的智能体对话生成交互式手势

Google Research推出AgentHands,一个由LLM驱动的XR原型,通过同步、富有表现力的手势增强对话智能体,在物理任务中提供空间锚定的指导,以弥合心理映射差距,增强用户参与度。该研究发表在CHI 2026上。

SOURCE / 全球热点解读 MIN / 9 ACCESS / 公开 POST / 2026-08-26 03:10:59

原贴

查看原文
作者:Google Research Blog 来源站点:research.google 原贴时间:

原文

Xun Qian, Research Scientist, and Ruofei Du, Interactive Perception & Graphics Lead, Google XR AgentHands is an LLM-powered XR prototype that augments conversational agents with synchronized, expressive hand gestures to provide spatially grounded guidance, bridging the mental mapping gap and enhancing user engagement in physical tasks. As AI assistants evolve from simple text interfaces to multimodal companions, we are seeing a shift toward more proactive, situated assistance. Recent innovations like Project Astra and Gemini 3.1 Flash Live already allow users to discuss their physical surroundings in real time, often utilizing visual bounding box overlays to identify objects in a camera feed. While these overlays are highly effective for 2D screens, the transition to immersive platforms like Android XR presents a unique challenge: how do we move beyond flat UI to create a truly embodied, spatially aware dialogue? To bridge this gap, we introduce AgentHands , published at CHI 2026 , a research prototype that brings the power of co-speech gestures to the 3D world. In human communication, our hands do more than just point; they describe shapes, mimic actions, and emphasize points, all synchronized with our voice. By leveraging the spatial understanding capabilities of Extended Reality (XR), AgentHands replicates this natural synergy. Following up our prior research in Human I/O and Sensible Agent , AgentHands further equips AI agents with expressive, synchronized hand gestures that transform abstract verbal instructions into intuitive, physical demonstrations, making conversations about your surroundings more natural and engaging. AgentHands demo: Empowering AI agents with expressive, synchronized hand gestures for spatially grounded conversations in XR. To start, we conducted a formative study with XR and human–computer interaction (HCI) experts at Google to determine what makes a virtual hand “legible” in a 3D environment. We distilled these insights into a multi-dimensional taxonomy that defines how an agent should use its hands to ground a conversation within a user's physical space. Handedness & gesture: Choosing between one or two hands and selecting from a library of forms, such as a “palm” for caution or a “cylindrical grip” to mimic holding a tool. Spatiality: Leveraging the depth of XR to determine where hands should live. They can be mid-air for general conversation, object-anchored for identifying specific parts, or user-relative for social cues. Temporal dynamics & visual effects (VFX): Gestures in XR aren't just static poses; they include animated motions like “pouring” or “tracing”. We also utilize XR’s unique visual layer by adding effects, such as a red glow to signify a heat warning. The AgentHands taxonomy diagram shows the six dimensions: Handedness, Gesture, Spatiality, Temporal Dynamics, Interactivity, and Visual Effects. The core innovation of AgentHands is its ability to map the high-level reasoning of LLMs into precise, real-time physical motions that match the agent's “voice” and the user's XR environment. We introduce the following key steps to compose the AgentHands workflow. The system begins with a lightweight object registration module. Using eye gaze and scene reconstruction, users can quickly “tag” items — like an orchid or a laptop — creating a spatial registry with 3D bounding boxes that the agent can reference.

中文翻译

Xun Qian,研究科学家,和Ruofei Du,交互感知与图形负责人,Google XR

AgentHands是一个由LLM驱动的XR原型,通过同步、富有表现力的手势增强对话智能体,以提供空间锚定的指导,弥合心理映射差距,并增强用户在物理任务中的参与度。

随着AI助手从简单的文本界面演变为多模态伴侣,我们看到正转向更主动、更具情境化的辅助。

像Project Astra和Gemini 3.1 Flash Live这样的最新创新已经允许用户实时讨论他们的物理环境,通常利用视觉边界框覆盖来识别相机画面中的物体。

虽然这些覆盖层在2D屏幕上非常有效,但向Android XR等沉浸式平台的过渡带来了独特的挑战:我们如何超越平面UI,创造真正具身化、空间感知的对话?

为了弥合这一差距,我们推出了AgentHands,发表在CHI 2026上,这是一个研究原型,将言语同步手势的力量带入3D世界。

在人类交流中,我们的手不仅仅是指示;它们描述形状、模仿动作、强调要点,所有这些都与我们的声音同步。

通过利用扩展现实(XR)的空间理解能力,AgentHands复制了这种自然的协同作用。

继我们之前在Human I/O和Sensible Agent方面的研究之后,AgentHands进一步为AI智能体配备富有表现力、同步的手势,将抽象的口头指令转化为直观的物理演示,使关于周围环境的对话更加自然和引人入胜。

AgentHands演示:为AI智能体赋予富有表现力、同步的手势,以在XR中进行空间锚定的对话。

首先,我们与Google的XR和人机交互(HCI)专家进行了一项形成性研究,以确定是什么让虚拟手在3D环境中“清晰可读”。

我们将这些见解提炼为一个多维分类法,定义了智能体应如何使用其双手来在用户物理空间中锚定对话。

手性&手势:选择一只手或两只手,并从形式库中选择,例如“手掌”表示谨慎,或“圆柱形握持”来模拟握住工具。

空间性:利用XR的深度来确定手应该放置的位置。它们可以在半空中用于一般对话,锚定在物体上用于识别特定部分,或相对于用户用于社交线索。

时间动态&视觉效果(VFX):XR中的手势不仅仅是静态姿势;它们包括动态动作,如“倾倒”或“描绘”。我们还利用XR独特的视觉层,添加效果,例如红色光晕表示热警告。

AgentHands分类图显示了六个维度:手性、手势、空间性、时间动态、交互性和视觉效果。

AgentHands的核心创新在于它能够将LLM的高级推理映射为精确、实时的物理运动,与智能体的“声音”和用户的XR环境相匹配。

我们引入了以下关键步骤来构成AgentHands工作流。系统从一个轻量级物体注册模块开始。利用注视和场景重建,用户可以快速“标记”物品——比如兰花或笔记本电脑——创建一个带有3D边界框的空间注册表,智能体可以引用这些边界框。

来源:Google Research Blog

核心信息

Google Research推出AgentHands,一个由LLM驱动的XR原型,通过同步、富有表现力的手势增强对话智能体,在物理任务中提供空间锚定的指导,以弥合心理映射差距,增强用户参与度。该研究发表在CHI 2026上。

  • Google Research推出AgentHands,一个由LLM驱动的XR原型,通过同步、富有表现力的手势增强对话智能体,在物理任务中提供空间锚定的指导,以弥合心理映射差距,增强用户参与度。该研究发表在CHI 2026上。
  • 原贴提到:Xun Qian, Research Scientist, and Ruofei Du, Interactive Perception & Gr
  • 来源:research.google

详细解读

这是什么信号?AgentHands是Google Research在CHI 2026上发表的研究原型,它揭示了AI交互正在从二维屏幕上的视觉叠加(如边界框)转向三维空间中的具身化表达。随着Android XR等沉浸式平台的发展,AI助手需要学会用“身体”说话,手势正是最自然的沟通方式之一。

为什么重要?当前的AI助手已经能通过语音和视觉识别物理环境,但输出方式仍局限于屏幕文字或2D标记。AgentHands提出了一套多维分类法(手性、手势、空间性、时间动态、交互性和视觉效果),为设计空间中的虚拟手势提供了系统框架。更重要的是,它证明了LLM的高级推理可以映射为实时的物理动作,这为未来AI助手在XR中提供直观指导奠定了基础。

对谁有价值?对XR开发者:可借鉴其分类法设计更自然的手势交互;对HCI研究者:提供了可验证的手势设计维度;对AI产品经理:启发了多模态AI助手的交互形态,尤其是在教育、维修、医疗等需要空间指导的场景。

可以怎么行动?可以深入研究CHI 2026论文,了解分类法细节;在自己的XR原型中尝试实现类似的手势系统;针对特定领域(如设备组装、烹饪指导)进行用户测试,验证手势辅助的实际效果。

风险或限制目前仍是研究原型,距离产品化尚有距离。手势的通用性和文化差异可能限制其适用性,且实时性能对硬件有较高要求。此外,用户对手势的接受度和学习成本需要进一步评估。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 research.google 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《AgentHands:在XR中为空间锚定的智能体对话生成交互式手势》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

AgentHands:在XR中为空间锚定的智能体对话生成交互式手势主要讲什么?

Google Research推出AgentHands,一个由LLM驱动的XR原型,通过同步、富有表现力的手势增强对话智能体,在物理任务中提供空间锚定的指导,以弥合心理映射差距,增强用户参与度。该研究发表在CHI 2026上。

这篇文章最值得关注的要点是什么?

Google Research推出AgentHands,一个由LLM驱动的XR原型,通过同步、富有表现力的手势增强对话智能体,在物理任务中提供空间锚定的指导,以弥合心理映射差距,增强用户参与度。该研究发表在CHI 2026上。;原贴提到:Xun Qian, Research Scientist, and Ruofei Du, Interactive Perception & Gr;来源:research.google

这篇文章和哪些AI专题相关?

它适合放在Agent工作流、AI日报、AI工具专题里阅读。 关联原因:这篇内容命中「Agent、智能体」等主题信号。;这篇内容命中「热点解读」等主题信号。;这篇内容来自该专题长期覆盖的栏目。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI日报、每日AI日报、AI信号、热点解读、BuilderPulse这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 规则洞察仪表板正式可用 下一篇 俄罗斯利用ChatGPT开展秘密影响力运动,在西方推动亲克里姆林宫叙事