Abstract
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. While this phenomenon has been examined in unimodal language models, its manifestation in vision-language models (VLMs) remains largely unexplored, and the field lacks a purpose-built instrument to investigate it.
We posit that studying contextual entrainment in VLMs requires more than porting existing text-only benchmarks to the multimodal setting; it necessitates a taxonomically structured, dual-modality instrument whose conditions are constructed around the item at hand (the depicted image in the textual stream, the textual query in the visual stream).
We argue that the transition to VLMs is substantive rather than incremental, making entrainment a dual phenomenon driven independently by textual and visual context. It introduces a veracity distinction (context that is false of the depicted scene yet possible in the world) that has no counterpart in the unimodal, world-knowledge-only formulations of prior work.
To make this position concrete and actionable, we introduce ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes, i.e., the association of context with the item and its relationship to truth, split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions).
We do not claim to measure entrainment in any particular model; we provide the instrument, the taxonomy that motivates it, and the evaluation protocols it enables, so that the community can investigate the phenomenon rigorously. We will release the dataset and its documentation publicly.
Blogger's Review: ENTRAP-VL offers a fresh perspective on the complexities of contextual entrainment in vision-language models. This work not only fills a significant gap in current research but also lays the groundwork for future model evaluations, aiding in the deeper understanding of how models behave with complex inputs. Its release is poised to greatly advance research in related fields.