NeFut Logo NeFut
中 Admin Login

[CS.AI] HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #LLM

Language models are typically evaluated on curated benchmarks that often under‑represent the complexity of real‑world enterprise deployments. We introduce HARDEN, a constrained evolutionary search method that adapts the inputs of existing evaluation cases into more challenging variants while keeping the expected outputs fixed. HARDEN searches along generated domain‑specific complexity axes and enforces feasibility constraints such as preserving task semantics, realism, and execution validity. Experiments on FinQA, PubMedQA, and ContractNLI using three scales of the Qwen3.5 model (35B‑A3B, 122B‑A10B, 397B‑A17B) show that HARDEN reduces task‑model accuracy by an average of 22.7% and up to 49.9% relative to single‑pass baselines under the same feasibility checks. These results demonstrate that evolutionary search can produce substantially harder yet valid evaluation cases.

Review

Original Source: https://arxiv.org/abs/2609.30571

[h] Back to Home