Official implementation of MSNav, a zero-shot vision-and-language navigation framework combining dynamic map memory, spatial reasoning, and LLM-based action planning.
Project and repository led and maintained by Chenghao Liu.
Our paper has been accepted by ICASSP 2026 🎉. Read it on IEEE Xplore or arXiv.
MSNav brings together three complementary modules for long-horizon navigation:
- Memory: Maintains a topological map and selectively prunes historical nodes to retain useful navigation context.
- Spatial: Uses Qwen-Spatial (Qwen-Sp), fine-tuned from Qwen3-4B, to infer relevant objects and destination layouts.
- Decision: Combines visual observations, navigation history, map memory, and spatial cues for GPT-based action planning.
Note: This repository contains navigation, spatial inference, and evaluation code. Datasets, observation images, and fine-tuned checkpoints must be prepared separately.
MSNav/
├── GPT/
│ ├── api.py # Vision-language API client
│ └── one_stage_prompt_manager.py # Navigation prompts and spatial cues
├── Spatial/scripts/
│ ├── infer_instr_obj.py # Instruction-to-object inference
│ ├── infer_instr_sr.py # Destination spatial reasoning
│ └── eval_obj_metrics.py # Object extraction evaluation
├── vln/
│ ├── main_gpt.py # Navigation evaluation entry point
│ ├── gpt_agent.py # Navigation agent and map pruning
│ ├── env.py # Matterport3D navigation environment
│ └── parser.py # Command-line configuration
├── utils/ # Data loading and logging
├── scripts/run.sh # Experiment configuration reference
├── figs/placeholder_pruned.png # Placeholder for pruned observations
└── requirements.txt # Core Python dependencies
git clone https://github.com/MrCapricornLiu/MSNav.git
cd MSNav
pip install -r requirements.txt
pip install h5pyAdditional requirements:
- Navigation: Matterport3D Simulator with
MatterSimPython bindings installed in the same environment. - Spatial inference: PyTorch, Transformers, ModelScope Swift (
ms-swift), and their dependencies.
The pinned requirements.txt covers core navigation dependencies, not the simulator or spatial-model stack.
Prepare R2R connectivity graphs, Matterport3D scan data, processed navigation annotations, and RGB observations. The navigation code expects the following layout:
DATA_ROOT/
├── R2R/
│ ├── connectivity/
│ └── annotations/
└── Matterport3D/
└── v1_unzip_scans/
IMG_ROOT/
└── <scan_id>/<viewpoint_id>/<view_index>.jpg
Use --root_dir for DATA_ROOT, --img_root for IMG_ROOT, and --split for the processed annotation JSON. The example below uses a filename containing processed to select the processed-data loading branch.
GPT/api.py: Setgeneration_keyand the APIbase_url.vln/gpt_agent.py: Pointself.placeholder_image_datato the includedfigs/placeholder_pruned.png.Spatial/scripts/: Set model checkpoints and input/output paths before spatial inference.
Keep API credentials local; do not commit them to the repository.
From the repository root, run the following command after completing setup. This example evaluates one instruction with map pruning enabled:
python -m vln.main_gpt \
--root_dir /path/to/datasets \
--img_root /path/to/RGB_observations \
--split /path/to/MapGPT_72_scenes_processed_1.json \
--start 0 \
--end 1 \
--output_dir output/msnav \
--dataset r2r \
--batch_size 1 \
--llm gpt-4o \
--response_format json \
--max_action_len 22 \
--max_tokens 1000 \
--save_pred \
--enable_map_pruningImportant:
--startis inclusive and--endis exclusive. Supply an explicit, valid--endfor the current processed-data loader, and keep--batch_size 1.
Append these arguments to the navigation command, continuing the preceding line with \:
--extended_instruction \
--extended_instr_file /path/to/extended_instructions.jsonThe extended JSON must contain matching scan, path_id, and instruction fields, together with final_destination_spatial_relations.
Outputs are saved under --output_dir:
- Trajectories and per-instruction metrics:
preds/case_InstrID_*.json(requires--save_pred). - Aggregate navigation metrics:
logs/valid.txt.
| Argument | Purpose |
|---|---|
--enable_map_pruning |
Enable dynamic map pruning. |
--pruning_start_step |
Set the first step at which pruning is considered. |
--map_pruning_step_threshold |
Set the minimum age for candidate nodes. |
--pruning_keep_recent_steps |
Protect recently visited nodes. |
--pruning_max_nodes_per_step |
Limit the number of nodes pruned per step. |
--w_time, --w_degree, --w_frontier |
Weight node age, connectivity, and unexplored neighbors. |
--enable_graph_distance_pruning, --w_dist |
Enable and weight the graph-distance component. |
--log_pruning_scores |
Log candidate scores for inspection. |
See vln/parser.py for defaults and additional options. The API client currently uses a fixed temperature of 0; there is no --temperature command-line option.
The spatial scripts use Qwen3-4B and configurable local checkpoints to extract objects and infer destination layouts.
| Script | Purpose |
|---|---|
infer_instr_obj.py |
Extract direct_obj and potential_obj lists from navigation instructions. |
infer_instr_sr.py |
Generate destination descriptions and spatial-layout cues. |
eval_obj_metrics.py |
Evaluate object extraction with F1, NDCG, and weighted metrics. |
After configuring the checkpoint and data paths in the inference scripts:
python Spatial/scripts/infer_instr_obj.py
python Spatial/scripts/infer_instr_sr.pyThe evaluation script retains local experiment configuration; review its model-loading and output-path settings before use. Qwen-Sp weights and I-O-S data are not bundled with the code.
If you use MSNav in your research, please cite our paper:
@inproceedings{liu2026msnav,
title = {{MSNav}: Zero-Shot Vision-and-Language Navigation with Dynamic Memory and {LLM} Spatial Reasoning},
author = {Liu, Chenghao and Zhou, Zhimu and Zhang, Jiachen and Zhang, Minghao and Huang, Songfang and Duan, Huiling},
booktitle = {2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
year = {2026},
url = {https://ieeexplore.ieee.org/abstract/document/11463005}
}This project is licensed under the MIT License.
We thank the authors of NavGPT, MapGPT, and InstructNav for their pioneering work in language-guided navigation. Their research and open-source contributions provide valuable foundations and inspiration for MSNav. We sincerely appreciate their efforts to advance the community.