Muon combines current and past gradients into a matrix momentum. For $M=U\Sigma V^\top$, the ideal polar update $Q=UV^\top$ gives every singular direction the same weight, which we call the flat profile. Recent optimizers replace this flat profile with fine‑grained spectral maps that assign each direction its own gain. This work asks how much spectral detail a Muon update actually needs.
Spectral diagnostics reveal that about $94\%$–$97\%$ of measured singular modes lie below an estimated noise edge, yet they collectively align positively with a reference gradient. We introduce BulkBoost, a two‑band spectral reweighting framework with fixed‑rank and noise‑calibrated variants. The latter uses split‑minibatch gradient differences to estimate a Marchenko–Pastur edge for Muon’s Nesterov input, separating the bulk below the edge from the spikes above it. Both variants increase the bulk’s relative weight through a shared gain while preserving the Frobenius norm of each unreweighted direction.
For a fixed partition, our theory provides the first‑order condition under which moving weight toward the bulk lowers the loss, and quantifies the fraction of the maximal first‑order improvement rate that two bands can capture. Across 30 continued‑pretraining settings (models from 14M to 410M parameters, six corpora), two‑band reweighting matches the fine‑grained power‑law profile of Freon and outperforms Spectra. Compared to Muon’s flat profile, Freon reduces final loss by an average of $0.022\%$ of the pre‑adaptation loss, while the two‑band variants achieve reductions of $0.073\%$–$0.147\%$. These observations suggest that useful departures from the flat profile are surprisingly low‑dimensional: a single bulk‑to‑spike gain captures at least as much benefit as the fine‑grained spectral profiles.
Review