Large language models (LLMs) have rapidly improved on long‑context tasks, largely thanks to Chain‑of‑Thought (CoT) reasoning. The exact way CoT helps a model count items scattered across a long text, however, remains unclear. To probe this, we created a needle‑in‑a‑haystack (NIAH) counting task: the model must report how many records appear in a lengthy document.\ \ Across twelve model comparison groups, models that performed explicit Thinking steps achieved higher counting accuracy than Non‑thinking models, with the gap widening for larger counts. This motivated a mechanistic analysis that identified two contrasting retrieval strategies:\ \
- Broad retrieval – Non‑thinking models attend to many potential needles simultaneously, spreading their attention.\
- Targeted retrieval – Thinking models enumerate needles step‑by‑step in the CoT trace, retrieving one needle at a time and concentrating attention on a single target.\ \ Targeted retrieval is accompanied by more compact internal representations: the hidden state becomes a tighter vector after each retrieval. Causal intervention experiments further suggest that Thinking models use the CoT trace to maintain and update an internal counter as needles are retrieved, even without explicit numbering; the counter state increments automatically.\ \ In small controlled experiments, both retrieval mechanisms and counter states emerge naturally under standard autoregressive training. Together, the findings link long‑context retrieval with the geometry of counting representations, supporting a state‑tracking account of CoT reasoning.\ \ Review: This work combines careful experiments with causal analysis to reveal that CoT improves long‑context counting by enabling targeted retrieval and compact representations, offering a clear mechanistic insight into how LLMs handle extensive context.