NeFut Logo NeFut
中 Admin Login

[CS.AI] Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Diversity collapse in parallel chain‑of‑thought has motivated inference‑time interventions that rely on a process reward model (PRM). The natural design is to prune a chain when the PRM deems it low‑scoring, extract its high‑PRM prefix, and graft it verbatim as an in‑context demonstration into a sibling that is still decoding. We isolate this mechanism as PRM‑Pruned Fragment Grafting (PPFG), the most cost‑minimal operationalization of cross‑trajectory step‑level transfer, and evaluate it at the operating point where earlier fragment‑grafting work only reported gains with additional compensating ingredients. Using Qwen2.5‑7B‑Instruct together with Math‑Shepherd on the full MATH500 benchmark (n=500, three random seeds), we compare both stagnation‑targeting and random‑targeting PPFG variants against an independent parallel‑CoT baseline. Across every measured axis the results are statistically indistinguishable. A four‑bucket classification of 322 stagnation‑rule injection events shows that only 14% actually target a genuinely struggling chain; the remainder hit chains that have already succeeded, are near completion, or sit on a flat PRM plateau—states that a rescue graft cannot alter. No compound‑gate refinement simultaneously achieves well‑targeted firing and sufficient density, and a random control matches the same parity at 2.4× the firing rate, indicating the inertness is not heuristic‑specific. The finding replicates across three base LMs, six benchmarks, a second PRM, and a compatibility‑gate sweep; two one‑sided tests promote the parity to positive equivalence on all twelve Qwen/LLaMA cells. Spot‑checking per event reveals injected chains prune at 2.75× the matched‑step rate, yet a surviving‑sibling counterfactual shows no population‑level compensation. A hindsight oracle bounds any per‑problem gain from choosing PPFG over independent at +0.13 percentage points. We also contribute an equivalence‑testing template for establishing inference‑time mechanism nulls, with every claim scoped to its tested operating point.

Review

Original Source: https://arxiv.org/abs/2610.00047

[h] Back to Home