Zero-shot vision-language navigation

NavJev: Efficient Vision-Language Navigation

via Action-Centric Visual Compression and Discriminative Action-Semantic Memory

Kai ShengLiuyi Wang†Jinlong LiHaojie DaiChengju LiuQijun Chen†

College of Electronic and Information Engineering, Tongji University, Shanghai, China · †Corresponding authors

{2610859, wly, li_jinlong, tju_dhj, liuchengju, qjchen}@tongji.edu.cn

Paper Code · coming soon View results ↓
27.0%Success rate
22.4%SPL
$0.06per 100 episodes
4.46 GiBpeak local GPU memory

Abstract

Efficient zero-shot VLN.

Recent zero-shot Vision-and-Language Navigation (VLN) methods increasingly rely on multimodal large language models (MLLMs) to reason over visual observations, navigation instructions, and candidate actions. Although effective, repeatedly invoking autoregressive multimodal reasoning at every navigation step introduces substantial inference latency, limiting the responsiveness of embodied agents. We propose NavJev, an efficient VLN framework that reformulates online navigation from repeated multimodal generation into compact visual compression followed by lightweight typed action selection. Action-Centric Visual Compression (ACVC) integrates waypoint geometry, BLIP captions, and RAM semantic tags into compact representations of candidate actions, while Discriminative Action-Semantic Memory (DASM) filters shared semantics and maintains discriminative action-specific evidence across navigation steps. Based on these representations, Jev directly performs structured probabilistic decisions over the available action set. Experiments on R2R-CE show that NavJev achieves 27.0% SR and 22.4% SPL with only 0.65 s per navigation step, while substantially reducing inference latency and cost compared with MLLM-based VLN methods.

NavJev performance, average step time, and API cost compared with representative VLN methods

Method

NavJev pipeline from panoramic observation and waypoint prediction through ACVC, DASM, and Jev typed navigation decisions
01 · ACVC

Action-Centric Visual Compression

Each executable waypoint is represented by its motion, heading, distance, BLIP caption, and RAM tags, converting panoramic RGB-D observations into compact candidate action states.

02 · DASM

Discriminative Action-Semantic Memory

Shared scene tags are removed while the selected action and its distinctive semantic evidence are retained across steps, preserving what separates one route choice from another.

03 · JEV

Typed Navigation Decisions

Jev receives the instruction, compact action options, and navigation memory, then returns one constrained option label that maps directly to an executable waypoint or STOP.

Evaluation

Demos

Simulation

Episode 259SR 1 · SPL 1.000
Episode 218SR 1 · SPL 1.000

Real World

Real-world NavJev trajectories from the paper in cafe and office environments

Citation

Cite NavJev

@misc{sheng2026navjevefficientvisionlanguagenavigation,
  title         = {NavJev: Efficient Vision-Language Navigation via Action-Centric Visual Compression and Discriminative Action-Semantic Memory},
  author        = {Kai Sheng and Liuyi Wang and Jinlong Li and Haojie Dai and Chengju Liu and Qijun Chen},
  year          = {2026},
  eprint        = {2609.34969},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2609.34969}
}