Multimodal large language models commonly reuse visual encoders pretrained with CLIP, even though the features of these vision transformers are ultimately consumed by autoregressive LLMs. The paper names this mismatch the semantic-interface gap.

MedMLIP pretrains the visual encoder through report generation with a frozen LLM — the encoder learns features useful for the LLM's actual job — while Local Relational Distillation preserves relationships among visual patches to avoid visual collapse.

Pretrained on IU-Xray and Open-PMC-300K and evaluated on VQA-RAD and SLAKE, only the ViT transfers while the guiding LLM and projector are replaced, which lets cross-LLM transferability be assessed directly. Code and the pretrained model were released on GitHub.

CLIP-style contrastive pretraining optimizes for retrieval-style alignment, not for feeding an autoregressive decoder.

Report generation with a frozen LLM ties the encoder's objective to the interface it will actually serve.

Distilling patch relationships keeps fine-grained visual information from collapsing under the language objective.

Replacing the downstream LLM at evaluation isolates what the encoder itself learned.

When a component's output is consumed by another model, align its training with that consumer's interface; generic proxies leave a gap the consumer pays for.

The cross-LLM transfer experiments demonstrated the value of pretraining visual encoders for their autoregressive LLM interface; code and the pretrained model are available on GitHub (SkyCol/MedMLIP).

FOLLOW THE EVIDENCE

The sources

  1. Pretraining of Medical Visual Encoders Toward Multi-modal Large Language Models arxiv.org