We train a model on the same set of problems using two incompatible yet correct conventions and ask how the ordering of the data is written into the parameters. The learning‑rate schedule is no longer a background condition; it acts as the averaging operator that decides the answer. We prove a bound where arrangement and schedule appear as separate multiplicative factors:
$$E = B \cdot w$$
Here $B$ denotes the arrangement’s block period and $w$ the weight the schedule can place on any single moment of the run. A decaying schedule cannot simultaneously have a large step size and an uncontracted remainder at the same moment, whereas a constant schedule does exactly that at the final step.
In the experiments we fix the corpus, budget, and training path, varying only the learning‑rate scheduler type. Two families are compared:
- Constant rate: the interior allocation spans $0.2221$, contrast floor $11.63$, and varies monotonically with how blocked the arrangement is.
- Single cosine (the default in all published arms): the same ten arrangements occupy two distinguishable states, even though their resolution would theoretically allow about ten.
Thus “order matters” and “order does not matter” are the two extremes of a single knob. What the path writes determines which convention the model commits to, and no exact‑match benchmark can detect it. Across twelve arms the sum $\text{acc}_A+\text{acc}_B$ stays within $9.7\%$, while the allocation share ranges from $0.04$ to $0.87$, meaning the $12.29\sigma$ arrangement switch measured in this paper is exactly zero under a convention‑agnostic metric. Marking the convention explicitly in the prompt collapses the switch, reaching $87.5\%$ of the union ceiling.
Review: By separating the effects of data arrangement and learning‑rate schedule, the paper provides a clear theoretical framework and controlled measurements that illuminate how models commit to conventions, offering a valuable tool for future studies of order sensitivity.