The "Linear Representation Hypothesis" (LRH) has been mentioned across AI, neuroscience, and cognitive science, but prior work has not consistently treated it as a falsifiable scientific claim. We first map the sources of inconsistency, focusing on the model under study, the location of the representation, the definition of features, and the evaluation dataset. We argue that only when these dependencies are made explicit can statements about linear representations be meaningfully tested. To this end, we propose a rigorous formalization that incorporates model, representation layer, feature definition, and dataset as explicit parameters, allowing LRH to be evaluated as a falsifiable hypothesis. Finally, we highlight several non‑trivial open problems, such as cross‑model representation consistency, generalization across task datasets, and quantitative neural interpretability, urging the community to address them.
Review