Long‑context understanding demands that large language models reason over documents, conversations, or code that span tens of thousands of tokens, yet task‑relevant evidence is often sparse and scattered among abundant irrelevant or redundant material. To address this, we propose the Highlight‑Then‑Summarize (H2S) paradigm: first locate source‑grounded, question‑relevant evidence in the raw text, then compress those evidence pieces into a concise, question‑conditioned summary, and finally generate the answer based on the summary.
To train this behavior we build the H2S‑Dataset, containing 6,647 examples from 11 benchmark families with an average context length of 43.9K tokens, and introduce H2S‑RL, which supplies process‑level rewards for evidence selection and summary construction in addition to the final‑answer correctness reward.
We evaluate on H2S‑Bench, a suite of seven long‑context tasks. Under a shared 128K input and 4K output budget, H2S‑14B achieves an average score of 32.60, surpassing Qwen3.8‑27B by 10.17 points and attaining the strongest overall result among the open‑source models tested. H2S‑14B also attains the highest Evidence‑Summary Quality score and retains 97.1% of its 16K‑budget performance when limited to a 4K output budget. These findings demonstrate that explicit evidence selection and integration improve long‑context reasoning while enabling more compact generation.
Review