One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

arXiv:2608.175122026-08-17

Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce a Pixel-to-3D Action Formulation (Point) tha

Authors · Hongyan Feng

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB