Long‑context sequence models face a fundamental trade‑off: softmax attention offers flexible token‑level interactions at a quadratic cost $O(T^2)$, while linear attention achieves $O(T)$ training and constant‑time decoding by compressing history into a fixed‑size state.\ \ This work introduces Structured Matrix Attention (SMat‑Attention), a family of causal masks whose row supports have VC‑dimension $d$. When $d=1$ the mask reduces to the standard causal mask; increasing $d$ enables richer subset‑routing patterns.\ \ To attain hardware efficiency, chunkwise forward and backward algorithms are proposed. For a sequence of length $T$, the hard‑routing construction requires $$O\bigl(T^{2-3/d}+T\bigr)$$ work, despite the mask being dense, thus preserving sub‑quadratic complexity.\ \ In a fixed‑horizon streaming scenario, decoding after a distant prefix takes constant time per token, with cached state size $$O\bigl(T^{1-1/d}\bigr)$$. Consequently, the VC‑dimension $d$ becomes an explicit knob governing access‑pattern complexity, prefill cost, and decoding memory.\ \ Empirical studies on subset‑routing and rule‑assisted multi‑key retrieval demonstrate the expressive power of the masks. Extensions to Mamba‑2 and Gated DeltaNet using learned routing with top‑$k$ query reads retain sub‑quadratic prefill, improve recall accuracy over the backbones in several settings, and achieve comparable performance on small‑scale language modeling tasks.\ \ Review: SMat‑Attention bridges the flexibility of soft attention with the efficiency of linear attention by introducing a tunable VC‑dimension. The approach opens a new design space for long‑context modeling, and experiments confirm its dual benefits in speed and accuracy.