摘要:端侧具身智能将具身大模型从云端远程推理推向设备本体的闭环执行,是机器人进入真实物理世界的关键系统基础。揭示视觉-语言-动作模型(VLA)与世界-动作模型(WAM)两类模型在多模态感知、扩散动作生成及世界模型滚动中产生的高计算与高访存开销,与端侧硬件在10~100 Hz控制频率、有限功耗及网络不稳定条件下可靠运行之间的系统性矛盾。针对该矛盾,提出以动作成功率、闭环时延与能效为共同优化目标,围绕少步化生成、动作感知压缩、缓存复用、异步分块执行与专用运行时等策略,开展算法与系统层面的协同设计。
关键词:端侧智能;具身智能;端侧推理优化
Abstract: Endowing embodied intelligence on the device side shifts large embodied models from cloud-based remote inference to on-device closed-loop execution, which serves as a crucial system foundation for robots to operate in real physical environments. The systematic contradiction between the high computational and memory-access overheads introduced by multimodal perception, diffusion-based action generation, and world model rolling in vision-language-action models (VLA) and world-action models (WAM), and the stringent requirements for reliable operation on end-side hardware under 10–100 Hz control frequencies, limited power budgets, and unstable network conditions, is revealed. With action success rate, closed-loop latency, and energy efficiency as the common optimization objectives, algorithm-system co-design strategies are proposed, centering on few-step generation, action-aware compression, cache reuse, asynchronous chunk execution, and specialized runtimes.
Keywords: on-device intelligence; embodied intelligence; on-device inference optimizations