Existing bias auditing methods usually rely on model outputs, requiring costly benchmarks or judge models and often missing shifts that only appear inside internal representations. We introduce a reference‑based auditing framework that directly compares hidden‑state representations across model variants such as before and after fine‑tuning. Because fine‑tuning reshapes the geometry of representations, absolute hidden vectors are not comparable. We first build a fixed set of anchor sentences and encode each evaluated sentence as its similarity vector to these anchors, yielding relative representations in a shared comparison space. We then measure how target groups shift their association with positive and negative attributes, defining the Representational Bias Shift as $$\Delta B$$. Experiments on three model families and the WildGuardMix, DecodingTrust, and ToxiGen benchmarks show that $$\Delta B$$ correlates with output‑level bias change in 15 out of 18 settings, reaching $|r|=0.84$ ($p<0.01$), indicating that relative representations reliably capture hidden‑layer bias migration.
Review