매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

arXiv:2608.175122026-08-17

Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce a Pixel-to-3D Action Formulation (Point) tha

저자 · Hongyan Feng

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사