From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
There is growing interest in biomedical vision--language models trained on scientific literature. However, most pipelines compress rich multi-panel figures and long captions into coarse figure-level pairs, discarding the fine-grained correspondences clinicians rely on when zooming into local structu