NeFut Logo NeFut
Admin Login

[CS.AI] CriticGen: Generation-Aware Evaluation as Actionable Feedback

Published at: 2026-09-10 22:00 Last updated: 2026-09-12 06:35
#AI #Machine Learning #LLM

CriticGen is a fine‑grained, generation‑aware evaluation framework that turns evaluation into actionable answer improvement control. It first generates sample‑specific evaluation dimensions and scoring criteria under high‑level categories such as subjective, objective, and self‑derived constraints. These criteria form a dynamic rubric. Under the same rubric the model jointly produces a score, a reason, an executable refinement suggestion, and a refined answer. This enables the model to diagnose flaws and perform targeted improvements. Experiments show that fine‑grained evaluation must be instance‑specific and actionable. The rubrics induced by CriticGen raise relevance/coverage from 3.33/4.03 to 3.97/4.24. Correlation with human labels reaches 0.9556 Pearson and 0.9560 Spearman. Criterion‑grounded reasons and executable suggestions improve their F1 from 0.6369/0.5994 to 0.7554/0.7900. Crucially, the feedback translates into reliable answer improvement, enhancing 73.17% of answers with a 93.28% non‑degradation rate.

Review

Original Source: https://arxiv.org/abs/2609.05439

[h] Back to Home