We investigate how Vision Transformers ground abstract concepts (e.g., angry) when training data provide limited direct referential evidence. The central hypothesis is a metonymic grounding mechanism: abstract predictions are driven by concrete, interpretable anchor concepts (e.g., fire) that bridge visual signals to abstract semantics.\ \ To test this, we apply Transcoders to CLIP and DINO vision encoders, recovering intermediate features that can be associated with semantic labels of more concrete concepts, and we trace their contributions within the circuits underlying abstract concept recognition.\ \ Experiments on a carefully curated icon dataset reveal structured metonymic circuits: perceptual primitives dominate early layers, object‑like anchors appear next, and finally abstract targets are reached. Images containing rendered text follow a distinct perceptual‑to‑textual route.\ \ Causal interventions further validate that metonymic intermediates are functionally involved in grounding abstract concepts.\ \ Review