解法
Real-world document processing relies on rigid predefined schemas, yet critical target fields often have no direct visual counterpart on the page — extracting them requires multi-hop derivation such as aggregating sub-categories or reasoning over visual marks. Even after standard fine-tuning, models retrieved incorrect visual evidence, or retrieved it correctly and then skipped intermediate derivation steps.
DocMIDE trains compact vision-language models to retrieve visual evidence explicitly before deriving an answer. Generation is constrained to a plan-retrieve-derive structure, optimized with Group Relative Policy Optimization under a four-component rule-based reward covering output format, the retrieved evidence block, every intermediate step, and the final value against a verified reference trace.
On a 4,151-pair implicit extraction benchmark, DocMIDE raised accuracy on Qwen3.5-4B from 70.8% to 95.9% from only a small set of annotated examples, and transferred to a second backbone architecture. Supervised demonstrations alone did not close the gap at any budget tested; rewarding the intermediate steps is what did.
What it achieved
Reward the intermediate derivation steps, not just the final answer
为什么有效
Implicit fields require chaining evidence across the page, so failures concentrate in skipped or wrong intermediate steps.
Forcing a plan-retrieve-derive structure makes each hop explicit and checkable.
Rule-based rewards on verified reference traces avoid the cost of human-grading every step.
The step-level reward, not extra demonstrations, produced the 25-point gain.
可以借鉴什么
When a task hides multi-hop reasoning, supervise the steps with rule-based rewards; demonstrations alone teach answers, not derivation paths.
后续
The paper reports accuracy rising from 70.8% to 95.9% on Qwen3.5-4B, transfer to a second backbone architecture, and the negative result that supervised demonstrations at any tested budget could not match rewarding intermediate steps.
从这里,追溯依据