A small number of unusually high‑gain parameters can exert disproportionate effects in transformer language models. This work asks whether analogous structures recur in genomic foundation models and whether structural geometry determines functional importance. We examined high‑gain rows in gated feed‑forward networks (gated‑FFN) across text and genomic foundation models, employing a frozen 22‑model causal census and computing the associated bilinear weight operator exactly, without a diagonal approximation.
Activation‑derived candidate rows were functionally enriched compared with random and top‑norm same‑layer controls. However, neither spectral concentration nor operator magnitude predicted causal effect size, and these associations vanished within the endpoint‑homogeneous text‑decoder subset.
A within‑layer sweep of 36 rows in one genomic decoder and one text decoder revealed two regimes: below the detector’s acceptance threshold the ratio carried no positive information about causal damage; above it the ratio ordered rows strongly but did not grade severity as a dose‑response. The sweep also uncovered a second individually catastrophic row invisible to a one‑candidate‑per‑model census, and non‑additive damage among co‑located critical rows.
Case studies showed divergent causal organizations: a robust super‑additive pair interaction in DNABERT‑2, and in GENERator a sharply position‑localized dependence where preserving or restoring the row’s beginning‑of‑sequence contribution rescued essentially all native‑loss damage.
Thus, high‑gain gated‑FFN rows constitute a recurrent architectural phenotype whose structural prominence acts as an enrichment signal, not a calibrated measure of functional criticality or a specification of causal organization. Enrichment is general, but the mechanism is model‑specific.
Review: By combining exact operator analysis with layer‑wise sweeps, the paper uncovers both shared and model‑specific aspects of high‑gain FFN rows, offering a fresh perspective on the functional relevance of transformer architecture. Future work could explore leveraging these structural signals for model interpretability and safety assessments.