每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

arXiv:2608.175122026-08-17

Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce a Pixel-to-3D Action Formulation (Point) tha

作者 · Hongyan Feng

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道