In the evaluation of molecular encoders, predictive accuracy is often the sole metric; however, it does not establish that the learned representation captures reproducible scientific structure, adds information beyond strong baselines, or transfers out of distribution. This paper presents a reliability-aware audit of generic molecular representations for human olfaction across four distinct claims: global perceptual geometry, incremental predictive value beyond chemistry, cross-dataset replication, and mixture transfer to unseen components.
Using the Keller-Vosshall and Bierling single-molecule rating datasets and the Ma binary-mixture dataset, we compare MoLFormer and ChemBERTa against RDKit descriptors and Morgan fingerprints under identity-controlled and matched evaluations. Human three-attribute rating geometry, based on intensity, pleasantness, and familiarity, is reproducible across participant splits (median RSA 0.743 and 0.855), while model-human alignment is substantially weaker (RSA 0.019-0.158). Learned embeddings do not consistently outperform conventional representations in global alignment, and MoLFormer shows no clear incremental predictive value beyond a combined RDKit-Morgan baseline in either single-molecule dataset. Human geometry exhibits positive but incomplete agreement across 63 shared molecules (RSA 0.331; 95% bootstrap interval [0.204, 0.507]). Under one strict unseen-component mixture split, incremental effects are outcome- and representation-dependent, with all intervals crossing zero.
These results establish empirical boundaries for the evaluated generic molecular encoders and motivate a broader evaluation principle: representation quality in scientific domains should be assessed separately for target reliability, structural alignment, incremental information, replication, and out-of-distribution transfer.