Human knowledge is organized in a hierarchical, inter‑dependent manner: mastering a concept typically requires prior mastery of its prerequisites, a principle formally captured by Knowledge Space Theory (KST). Large language models (LLMs) have recently achieved strong performance on complex reasoning tasks, yet it remains unclear whether they possess a comparable, human‑like knowledge structure. To address this, we propose a KST‑based evaluation framework that examines whether LLMs’ mathematical reasoning respects the same prerequisite dependencies observed in human learners. We benchmark eight open‑ and closed‑source LLMs against real human participants and obtain three key findings:
- LLMs frequently violate knowledge dependencies and fail to exploit related information provided in the context to improve performance on downstream dependent questions;
- The knowledge structures of different LLMs are inconsistent, as evidenced by a low overlap in their knowledge distributions;
- These structural shortcomings are largely invisible to conventional accuracy‑based metrics and to evaluations where LLMs act as judges.
Taken together, our behavioral evidence indicates that the knowledge held by current LLMs does not follow a human‑like, structured organization.
Review