NeFut Logo NeFut
Admin Login

[CS.AI] LLMs as Post‑hoc Auditors of Physiological Plausibility in Symbolic Regression: A Clinician‑Evaluated Case Study

Published at: 2026-09-12 22:00 Last updated: 2026-09-15 01:15
#AI #Machine Learning #LLM

Genetic Programming and its variants such as grammatical evolution are widely employed in Symbolic Regression to derive mathematical expressions from multivariate data. Beyond predictive accuracy, researchers value the potential interpretability of these models, i.e., explicit equations linking inputs to outcomes. In practice, evolved models can become overly complex or conflict with established scientific knowledge, making interpretability and plausibility hard to achieve.

This study investigates whether Large Language Models (LLMs) can assist in enhancing the explainability of Symbolic Regression models generated by evolutionary computation. Building on our previous grammar‑based Genetic Programming work for estimating body‑fat percentage, we employ LLMs as post‑processing tools to analyze evolved expressions and rank them according to interpretability and medical plausibility. Four symbolic expressions were examined by three different LLMs, each run three times, and the resulting interpretations and rankings were evaluated by a panel of three clinicians.

Across the three LLMs, comparative model‑ranking outputs received more favorable assessments from clinicians than isolated term‑level explanations, suggesting that comparative audit information is more useful. However, the LLMs also produced physiologically and mathematically questionable explanations, indicating that they are better suited for comparative auditing under expert oversight rather than autonomous validation.

This paper is an extended version of a manuscript submitted to a journal.

Review

Original Source: https://arxiv.org/abs/2609.11431

[h] Back to Home