This paper investigates how semantic context shapes the geometry of learned vector representations. Conventional retrieval often relies on cosine similarity, which imposes a single fixed geometry. In contrast, semantic similarity is inherently context‑dependent: two images may be considered similar because they depict the same object, share a visual style, or correspond to the same clinical finding.
The key idea is to intertwine contrastive learning, exponential families, and information geometry, establishing a correspondence between probability distributions over "anchors" and Bregman geometries on the representation space. Under this correspondence, the anchor distribution itself determines the geometry.
Leveraging this insight, the authors introduce Anchor Divergence, a method for specifying context‑specific semantic geometries on fixed representations. Modeling the anchor distribution is thus equivalent to modeling the geometry of semantic similarity.
Experiments on image retrieval demonstrate that anchor divergences provide an effective and efficient way to encode context‑specific semantic similarity, leading to notable performance gains.
Review