IBBench-Light is designed to evaluate a model's responses to two kinds of external directives—apply and read—on the same record. Twelve semantic bases generate 144 matched pairs per model; four quantized instruction models produce a total of 1,152 greedy responses.
Paired Exact-Contract Accuracy (PECA) requires both members of a pair to satisfy their respective output contracts. Qwen succeeds on 132 execute prompts and 109 process prompts, yet only 97 complete pairs are correct, showing how marginal averages can hide important details.
We audit literal‑target exposure and case normalization, then add 1,722 logged CPU generations as directive‑absent controls, twelve additional semantic bases, intra‑base wording variations, and generation‑stopping conditions.
In the pinned Phi rerun, changing the end‑of‑sequence (EOS) set raises exact paired success from 0/144 to 62/144. A bounded IHEval comparison uses the same SmolLM2 checkpoint and output budget while preserving its published instruction roles and scorer. The benchmark measures conditional task and output‑contract success; task margins and paired counts must be interpreted together with the stopping policy.
Review: IBBench-Light’s paired approach reveals fine‑grained performance gaps across directives, highlighting the pivotal role of stopping policies and output contracts in model evaluation.