In this paper, we investigate whether vision foundation models construct representations that reflect the intrinsic properties of 3D Euclidean space. Unlike previous works that assessed 3D awareness of vision features by regressing image-centric quantities such as depth or normals, we focus on the relationship between the structure of the visual feature space and the group of Euclidean transformations $SE(3)$. We propose a set of probes to evaluate this relationship from both topological and geometric perspectives: a mutual neighborhood metric that measures the alignment between feature neighborhoods and spatial topology, and a Poincaré Adapter to test the linear accessibility of the geometry of camera motion from latent displacements in static scenes.
We demonstrate that self-supervised vision models, which have not been trained with direct 3D supervision or active agency, possess latent subspaces that are remarkably correlated with three-dimensional Euclidean space when probed correctly. Building on this insight, we propose a new class of "Latent-Space Navigation" techniques that perform visual odometry and localization purely in the latent space, bypassing the need for explicit 3D reconstruction.
Blogger's Review: This study offers a fresh perspective on the 3D spatial representation capabilities of self-supervised vision models, revealing the close relationship between latent space geometry and the real world. This finding could lead to the development of more efficient visual navigation algorithms, reducing reliance on traditional 3D reconstruction and holding significant practical implications.