CARAT fixes the question and gold answer while evaluating across eight matched views, and names each structural relation separately in GraphSpace. It adds matched fine‑tuning, answer masking, evidence injection, paired inference, and a rule that can withhold claims.
On the hardest families, the grounded view outperforms formula‑only inputs by 17.3 points. Further analysis shows that GraphSpace beats a plain periodic graph by 19.3 points, but this margin consists of two effects: when the plain rendering already contains everything the question needs, the gain is only 1.96 points; when it omits those fields entirely, the gain jumps to 46.7 points. In other words, the headline mainly measures what the baseline lacked rather than how evidence is presented.
We also turned the scrutiny onto our own benchmark. A rule that skips the link and reads the list directly answers four of seven hardened families, so we rebuilt it until eleven such shortcuts hovered near chance. The frozen model quotes the link yet gives the same answer when we redirect it in 95.6% of paired cases: it repeats the relation without using it. After matched supervision its accuracy reaches 99.8%, and deleting the link drops it to 23.4%, below the 27.0% achieved by the best shortcut. Both steps are therefore learnable.
Review