Dev.to | 📄 原文链接 | 2026-07-28 收录

视觉-语言-动作模型:当大语言模型学会使用双手

来源:dev.to — 2026-07-25

📋 概述

一只仿生手握住酒杯需要每 8-20 毫秒更新一次动作估计,否则要么捏碎杯子要么松手掉落。而视觉语言模型处理同一场景可以悠闲地花 100 毫秒无人察觉。NVIDIA GR00T N1.7、Google DeepMind Gemini Robotics 1.5 和 Physical Intelligence pi-zero-point-five 三款最新 VLA 模型同期发布,都自称通用操作机器人系统,却在思考和行动的缝合点上做出了不同的架构选择。

🔑 核心要点

💡 金句

A humanoid hand closing on a wine glass needs a new action estimate every 8 to 20 milliseconds, or it either crushes the glass or drops it.
← 返回 Dev.to 首页