Vision-Language Navigation (VLN) enables robots to follow natural‑language instructions for moving through an environment, making human‑robot interaction more intuitive. Existing VLN models rely on navigation graphs, panoramic views, and precise localization, which are hard to obtain in real settings. This work introduces a VLN approach that operates in continuous space without requiring graphs or panoramas and performs a simulation‑to‑real domain shift.
The core is a Cross‑Modal Attention (CMA) network that projects visual inputs and linguistic commands into a shared embedding space. The network is first pretrained on a public simulated dataset, then fine‑tuned with real‑world data collected by a custom Ackermann‑steered robot equipped with a forward‑facing camera and LiDAR. Captured images undergo linear photometric adjustment to narrow the sim‑real gap. Fine‑tuning uses only a few dozen episodes, after which the model runs offline on dedicated hardware in real environments.
Evaluation uses Success weighted by Path Length (SPL) and Normalized Dynamic Time Warping (nDTW). Results show that with limited real‑world fine‑tuning, the model’s SPL and nDTW in real scenes approach simulated performance, demonstrating robustness and adaptability. Review