This study investigates how a language model decides between two conflicting documents. We evaluated Google’s pre‑trained Gemma 4‑e4b on a targeted behavioral suite of 13 items, performing 784 forward passes in short, single‑turn contexts. A fully counterbalanced design allowed us to mathematically isolate the effects of source framing and reading position while cancelling the model’s inherent vocabulary bias.
The findings reveal that semantic framing of a source outweighs the order in which documents are presented. Labeling a document as an official guideline or a fresh update significantly shifts the model’s final answer. The model shows a primacy bias, preferring the first document, but the strength of this bias varies by at least a factor of five depending on subtle wording differences. Positional bias is driven by overall structural repetition rather than short trigger phrases; when the two documents share an identical template the primacy effect intensifies, whereas introducing wording variation reduces it.
Review