We present a six‑stage audit framework for assessing the reproducibility of scientific claims in the computer science literature, and instantiate it for the neuro‑symbolic AI (NSAI) subdomain.
Stage 1 retrieved 5,497 records; after removing 3,018 duplicates, 2,479 unique records remained.
Stage 2 screened titles and abstracts, identifying 1,365 self‑declared NSAI papers. Full‑text screening then excluded 61 more for being off‑topic, non‑research, lacking quantitative evaluation, or having inaccessible full text.
Stage 3 sought a verifiable public code artifact for each of the 1,304 eligible papers. No code artifact was found for 849 papers, leaving 455 to enter the artifact inventory and proceed to stages four and five.
In Stages 4 and 5, we fully or partially reproduced 85 studies, representing 6.52% of the eligible corpus and 18.68% of attempted reruns.
Reproduction attempts were blocked mainly by missing non‑code artifacts (321 cases) and missing or unusable code repositories (42 cases). These figures reveal a persistent reproducibility deficit that persists even when “code available” is claimed, underscoring the need for enforced, versioned, permanently archived artifact bundles in future NSAI publications.
Blogger's Review: This work quantifies the reproducibility crisis in neuro‑symbolic AI with a rigorous audit pipeline, offering concrete recommendations that could reshape publishing standards across the field.