Multimodal large language models sometimes receive conflicting signals from images or speech and the accompanying text. Existing measures of text bias often entangle modality preference with the order in which evidence is presented, making it unclear which factor drives the model's response. Prior studies either fixed the evidence order or moved task instructions together with the evidence, leaving the contribution of order ambiguous.
This paper adopts a paired‑comparison setup: the task instructions and the content of the two evidence sources remain unchanged, only their positions are swapped. This design isolates the effect of order on model judgments. Experiments on a range of vision and speech models show that placing an image or audio recording after conflicting text consistently pushes the answer toward the later perceptual input.
We also revisit earlier works and explain why their experimental configurations can lead to misleading conclusions. Overall, the findings reveal cross‑modal evidence noncommutativity: the same evidence can yield different judgments when its order changes, and positioning perceptual evidence later increases the model's reliance on its content.
Review