Large language models (LLMs) are widely evaluated on molecular property benchmarks, yet accuracy cannot tell whether a model predicts a property or simply retrieves a published number. We audited 22 frontier models across 12 regression benchmarks for verbatim retrieval and found the phenomenon to be common but highly benchmark‑specific: on five datasets more than $50\%$ of the models exhibit verbatim retrieval, while on the remaining datasets it appears only in isolated cells. Experiments were run at two reasoning levels, showing that reasoning influences retrieval; the same molecules and prompt flagged $89\%$ more often at the higher reasoning level than at the lowest. We then attempted to interrupt retrieval in the most contaminated cases and discovered that even the strongest models still recognize a combination of transformed SMILES strings and original labels. Suppressing retrieval brings the relative prediction errors of different models closer together, whereas the use of verbatim retrieval spreads them apart. This indicates that an LLM's overall predictive capability is not determined solely by the amount of memorised values. The paper provides an overview of the amount and depth of verbatim retrieval in molecular regression benchmarks using LLMs.
Review