解法
Multimodal large language models commonly reuse visual encoders pretrained with CLIP, even though the features of these vision transformers are ultimately consumed by autoregressive LLMs. The paper names this mismatch the semantic-interface gap.
MedMLIP pretrains the visual encoder through report generation with a frozen LLM — the encoder learns features useful for the LLM's actual job — while Local Relational Distillation preserves relationships among visual patches to avoid visual collapse.
Pretrained on IU-Xray and Open-PMC-300K and evaluated on VQA-RAD and SLAKE, only the ViT transfers while the guiding LLM and projector are replaced, which lets cross-LLM transferability be assessed directly. Code and the pretrained model were released on GitHub.
What it achieved
Pretrain the component for its real consumer, not a generic proxy
为什么有效
CLIP-style contrastive pretraining optimizes for retrieval-style alignment, not for feeding an autoregressive decoder.
Report generation with a frozen LLM ties the encoder's objective to the interface it will actually serve.
Distilling patch relationships keeps fine-grained visual information from collapsing under the language objective.
Replacing the downstream LLM at evaluation isolates what the encoder itself learned.
可以借鉴什么
When a component's output is consumed by another model, align its training with that consumer's interface; generic proxies leave a gap the consumer pays for.
后续
The cross-LLM transfer experiments demonstrated the value of pretraining visual encoders for their autoregressive LLM interface; code and the pretrained model are available on GitHub (SkyCol/MedMLIP).
从这里,追溯依据