2025
Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion
ICLR 2025poster
In computer vision tasks, features often come from diverse representations, domains (e.g., indoor and outdoor), and modalities (e.g., text, images, and videos). Effectively fusing these features is essential for robust performance, especially with the availability of powerful pre-trained models like…