Knowledge File / AI技能杠杆
论文速读:ReCAPA,解读最新 AI 进展
ReCAPA是一种预测对齐与规划架构,通过预测和对比在动作、子目标、轨迹三个层面调整偏差,使用Sinkhorn和Score-field模块强制语义对齐,减少错误传播。在VisualAgentBench、MineDojo、AI2-THOR等基准上超越强baseline。
SOURCE / AI技能杠杆
MIN / 4
ACCESS / 会员
POST / 2026-04-24 12:00:06
原贴
查看原文
原文
arXiv:2604.21232v1 Announce Type: new Abstract: Vision-Language-Action systems follow instructions to execute multi-step tasks in multimodal environments. Recent VLA approaches typically rely on post-hoc correction mechanisms or operate under fixed task decompositions and alignment schemes. However, once an intermediate step is mis-specified, local errors propagate through subsequent steps and eventually accumulate into cascading failures. To mitigate this compounding effect, we propose Predictive Alignment and Planning Architecture, a framework that uses prediction and contrast to adjust deviations across three levels: actions, subgoals, and trajectories. Semantic alignment is enforced at all levels using a Sinkhorn-based module and a Score-field module. The predictive correction and alignment jointly update the action generator during training, enabling it to adjust fine-grained steps to remain aligned with the overall intent. We further introduce two new metrics to quantify error propagation and recovery processes in tasks, capturing how mistakes spread and fade over long-horizon execution. Experiments show that ReCAPA achieves competitive results on embodied agent benchmarks such as VisualAgentBench, MineDojo, and AI2-THOR, outperforming strong proprietary and open-source Large Language Model baselines.
中文翻译
视觉-语言-动作系统在多模态环境中遵循指令执行多步骤任务。最近的VLA方法通常依赖于事后纠正机制或在固定的任务分解和对齐方案下运行。然而,一旦中间步骤被错误指定,局部错误会通过后续步骤传播,最终累积成级联故障。为了减轻这种复合效应,我们提出了预测对齐与规划架构,一种使用预测和对比在三个层面(动作、子目标和轨迹)调整偏差的框架。使用基于Sinkhorn的模块和Score-field模块在所有层面强制语义对齐。在训练期间,预测性纠正和对齐共同更新动作生成器,使其能够调整细粒度步骤以保持与整体意图一致。我们进一步引入了两个新指标来量化任务中的错误传播和恢复过程,捕捉错误如何在长程执行中传播和消退。实验表明,ReCAPA在具身智能体基准如VisualAgentBench、MineDojo和AI2-THOR上取得了有竞争力的结果,优于强大的专有和开源大语言模型基线。
核心信息
ReCAPA是一种预测对齐与规划架构,通过预测和对比在动作、子目标、轨迹三个层面调整偏差,使用Sinkhorn和Score-field模块强制语义对齐,减少错误传播。在VisualAgentBench、MineDojo、AI2-THOR等基准上超越强baseline。
- ReCAPA通过预测和对比纠正多步骤任务中的错误传播。
- 在动作、子目标、轨迹三个层面进行对齐。
- 使用Sinkhorn和Score-field模块强制语义对齐。
- 在VisualAgentBench等基准上优于开源和商业大模型。
- 提出了两个新指标量化错误传播和恢复。
试看内容
成为会员查看完整内容
你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。
详细解读
信息差价值
参考来源
成为会员查看完整内容