NeFut Logo NeFut
中 Admin Login

[CS.AI] DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

Real‑world document processing pipelines usually rely on rigid schemas, yet many target fields lack a direct visual counterpart and must be derived through multi‑hop reasoning, such as aggregating sub‑categories or interpreting visual marks. Existing approaches handle explicit text spans or single‑step implicit queries, but they fail at multi‑hop visual reasoning: models either retrieve wrong visual evidence or, even when retrieval is correct, skip the intermediate derivation steps.

To tackle this, we introduce DocMIDE, a fine‑tuning framework that trains compact vision‑language models to explicitly retrieve visual evidence before producing an answer. DocMIDE enforces a plan‑retrieve‑derive generation structure and optimizes it with Group Relative Policy Optimization under a rule‑based reward composed of four components—output format, retrieved evidence block, each intermediate derivation step, and the final value compared to a verified reference trace.

On a benchmark of 4,151 implicit extraction pairs, DocMIDE lifts Qwen3.5‑4B’s accuracy from 70.8% to 95.9% using only a small set of annotated examples, and the improvement transfers to a second backbone architecture. Pure supervised demonstrations cannot close this gap at any tested budget; rewarding the intermediate steps is the decisive factor.

Review

Original Source: https://arxiv.org/abs/2609.24092

[h] Back to Home