Synthetic perturbations are proposed as a cheap source of calibration data for LLM evaluators in biomedical machine learning, where expert review is limited. A planted mutation key, however, is neither a detector output nor an automatically generated human ground truth. We formalize four separate ledgers: planted perturbations, independent detector outputs, source‑linked human dispositions, and human‑added discoveries. We then audit the evaluation design, scoring code, read paths, and existing human records of a private synthetic Japanese care‑handoff workflow. The factory stored 69 planted error cards across 47 targets. Final review covered 22 targets and contained 22 confirmed imported proposals, 9 rejected proposals, and 79 human‑added cards; only 3 targets were double annotated. Passing the imported plant keys to a generic detector scorer yields 22/(22+9)=0.710 and 22/(22+79)=0.218. A direct audit of identity shows these numbers represent proposal‑confirmation yield and submitted‑ledger composition, not judge precision or recall, because no independent detector realization was retained for the audited proposals in the available records. The audit also uncovered source‑name collisions, row shadowing, forced severity, vacuous ratio defaults, and unsupported zero‑support field weights. We contribute a provenance‑aware claim audit, a storage contract, and a minimum calibration gate for responsibly communicating biomedical ML capability claims. This single‑workflow forensic case serves as an existence proof of a failure mode rather than an estimate of its prevalence: existing human work supports an exploratory audit of synthetic proposals, but not the operating characteristics of LLM judges, clinical validity, corpus prevalence, or robust inter‑annotator agreement.
Review