Real-world document processing relies on rigid predefined schemas, yet critical target fields often have no direct visual counterpart on the page — extracting them requires multi-hop derivation such as aggregating sub-categories or reasoning over visual marks. Even after standard fine-tuning, models retrieved incorrect visual evidence, or retrieved it correctly and then skipped intermediate derivation steps.

DocMIDE trains compact vision-language models to retrieve visual evidence explicitly before deriving an answer. Generation is constrained to a plan-retrieve-derive structure, optimized with Group Relative Policy Optimization under a four-component rule-based reward covering output format, the retrieved evidence block, every intermediate step, and the final value against a verified reference trace.

On a 4,151-pair implicit extraction benchmark, DocMIDE raised accuracy on Qwen3.5-4B from 70.8% to 95.9% from only a small set of annotated examples, and transferred to a second backbone architecture. Supervised demonstrations alone did not close the gap at any budget tested; rewarding the intermediate steps is what did.

Implicit fields require chaining evidence across the page, so failures concentrate in skipped or wrong intermediate steps.

Forcing a plan-retrieve-derive structure makes each hop explicit and checkable.

Rule-based rewards on verified reference traces avoid the cost of human-grading every step.

The step-level reward, not extra demonstrations, produced the 25-point gain.

When a task hides multi-hop reasoning, supervise the steps with rule-based rewards; demonstrations alone teach answers, not derivation paths.

The paper reports accuracy rising from 70.8% to 95.9% on Qwen3.5-4B, transfer to a second backbone architecture, and the negative result that supervised demonstrations at any tested budget could not match rewarding intermediate steps.

FOLLOW THE EVIDENCE

The sources

  1. DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents arxiv.org