Few-Shot Learning from Gigapixel Images via Hierarchical Vision-Language Alignment and Modeling
Vision-language models (VLMs) have recently been integrated into multiple instance learning (MIL) frameworks to address the challenge of few-shot, weakly supervised classification of whole slide images (WSIs). A key trend involves leveraging multi-scale information to better represent hierarchical t…