NeFut Logo NeFut
Admin Login

[CS.AI] TPvG: A Moral Decision Framework for Large Language Models from One-Shot to Sequential Feedback

Published at: 2026-09-02 22:00 Last updated: 2026-09-03 02:56
#AI #Machine Learning #LLM

Existing evaluations of LLM morality usually present models with isolated moral vignettes and ask for a single‑shot decision, overlooking a factor that profoundly shapes human moral behavior: consequence feedback. To address this gap we introduce TPvG (Text‑based Pain‑versus‑Gain), adapted from a human moral paradigm, which embeds consequence feedback into the everyday dilemma of not harming others versus maximizing self‑gain. TPvG comprises five moral decision tasks that progress from minimal‑context one‑shot choices to sequential decisions with explicit feedback about outcomes. Our experiments show that LLM moral choices are strongly influenced by the decision format (one‑shot vs. sequential), and that explicit receiver feedback yields heterogeneous effects across models. Moreover, LLM responses to explicit feedback diverge from the human reference pattern, suggesting potentially different decision processes. These findings highlight the need to assess whether LLM moral behavior remains stable in high‑stakes interactive settings.

Review

Original Source: https://arxiv.org/abs/2608.28610

[h] Back to Home