The solution
Real-world document processing relies on rigid predefined schemas, yet critical target fields often have no direct visual counterpart on the page — extracting them requires multi-hop derivation such as aggregating sub-categories or reasoning over visual marks. Even after standard fine-tuning, models retrieved incorrect visual evidence, or retrieved it correctly and then skipped intermediate derivation steps.
DocMIDE trains compact vision-language models to retrieve visual evidence explicitly before deriving an answer. Generation is constrained to a plan-retrieve-derive structure, optimized with Group Relative Policy Optimization under a four-component rule-based reward covering output format, the retrieved evidence block, every intermediate step, and the final value against a verified reference trace.
On a 4,151-pair implicit extraction benchmark, DocMIDE raised accuracy on Qwen3.5-4B from 70.8% to 95.9% from only a small set of annotated examples, and transferred to a second backbone architecture. Supervised demonstrations alone did not close the gap at any budget tested; rewarding the intermediate steps is what did.
Why it worked
Implicit fields require chaining evidence across the page, so failures concentrate in skipped or wrong intermediate steps.
Forcing a plan-retrieve-derive structure makes each hop explicit and checkable.
Rule-based rewards on verified reference traces avoid the cost of human-grading every step.
The step-level reward, not extra demonstrations, produced the 25-point gain.
What can be applied
When a task hides multi-hop reasoning, supervise the steps with rule-based rewards; demonstrations alone teach answers, not derivation paths.
Aftermath
The paper reports accuracy rising from 70.8% to 95.9% on Qwen3.5-4B, transfer to a second backbone architecture, and the negative result that supervised demonstrations at any tested budget could not match rewarding intermediate steps.
FOLLOW THE EVIDENCE