NavGen introduces a text‑to‑video pipeline powered by high‑fidelity visual generative models to create diverse vision‑language navigation (VLN) episodes for both indoor and outdoor environments. The pipeline consists of: 1) leveraging large‑scale pretrained diffusion or Transformer video generators to translate natural language commands into corresponding aerial videos; 2) applying a style‑diversification technique that randomly samples lighting, weather, material and other conditions to expand coverage of long‑tail scenarios. The resulting dataset contains roughly 400 K navigation episodes covering tasks such as path planning and goal localization. Empirical results show that model performance improves steadily with data scale and surpasses existing UAV navigation datasets across all metrics. Real‑world transfer is validated by deploying the model in a world‑action‑model framework on actual drones, achieving a 75% success rate.
Review