NeFut Logo NeFut
中 Admin Login

[CS.AI] Spatial Strategies, Not Actions: Vector-Quantized Geodesics as Tools for LLM-Driven Agents

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#algorithm #LLM #Artificial Intelligence

Large language model (LLM) based agents are often criticized for lacking spatial awareness and merely exploiting statistical text patterns. To probe their spatial comprehension, we built an architecture that couples geometric utilities with an LLM acting as a high‑level orchestrator in grid‑world environments.

Offline phase: The agent gathers a large set of geodesic trajectories in the environment. These trajectories are compressed via vector quantization (VQ) into a representative subset. An LLM then assigns a natural‑language description to each selected trajectory, turning it into a reusable "tool".

Online phase: Given the current state and goal, the LLM selects the appropriate tool; low‑level control is delegated to primitive actions that execute the trajectory associated with that tool.

From an agentic AI perspective, learning is split into two levels:

We evaluated the approach in a partially observable, dynamic 2D grid environment using the open vision‑language model Qwen3.6-35B-A3B. By pairing the geometry‑derived tool library with a zoom‑in tool and a collision‑detection tool, a fast, non‑reasoning configuration matched the goal‑reaching rate of a far more expensive chain‑of‑thought version, while reducing decision latency from minutes to seconds.

Results demonstrate that the offline cost of building the tool library is amortized, and online inference remains extremely cheap, validating the practicality of spatial‑strategy‑driven agents.

Review

Original Source: https://arxiv.org/abs/2610.00613

[h] Back to Home