The quadratic cost of dense self‑attention is a core bottleneck for long‑context language modeling. Existing efficient alternatives fix sparsity or locality before training, but natural‑language dependencies vary with each input and cannot be prescribed in advance. We treat attention approximation as a geometric problem and let the model learn the interaction geometry from data.
MoSAR (Mixture of Semantic Attention Regimes) introduces input‑conditioned query and key routers placed after positional encoding. The routers select a mixture of short, medium, and global regimes, yielding a continuous distance‑dependent attention field instead of a static sparsity pattern. This geometry is learned during training and can be discretized by top‑1 routing.
Controlled pre‑training with matched 500 M‑parameter models shows that MoSAR learns a substantially shorter‑reach attention geometry without hurting language‑model quality, achieving lower perplexity than dense RoPE at the training context length. Under length extrapolation, MoSAR attains the best perplexity among all baselines, including strong ones like ALiBi, and the learned geometry remains stable after deterministic top‑1 discretization, indicating both adaptivity and suitability for low‑cost inference.
Review