Large vocabularies impose a heavy inference cost on the output heads of small language models. We introduce a post‑training technique called softmax reparameterization that searches over functionally equivalent output heads before quantization. The method subtracts a scalar multiple of the vocabulary‑row mean from each output row and selects the coefficient by minimizing validation KL divergence. For linear‑softmax heads this shift preserves full‑precision predictions exactly, requiring no decoder retraining; a rank‑one correction extends the approach to nonlinear logit paths. Experiments span seven heads and three quantizers. W4 gains are largest where baseline quantization severely distorts predictions: test KL drops by 93% on XGLM with RTN and by 73%‑77% on Phi, BLOOM and BLOOMZ with activation‑weighted MSE. Heads with low baseline error change little; under W2 compression stress the benefits become more widespread. On Phi the improvements survive stronger GPTQ calibration, and an untouched holdout reproduces the gains on Phi and BLOOM. Coefficients selected on WikiText transfer to C4 and OpenWebMath without retuning. Residual analysis on Phi shows that fidelity can improve despite higher total logit error: the chosen representative reduces error on likely outputs and lowers its Fisher‑weighted cost. For shift‑compatible heads the shift adds no inference operation. With the decoder kept in BF16, a packed W4 Phi output head reduces batch‑one generation latency by 10.8%, and reparameterization preserves this speedup.
Review