2025
MixSense : Mixture of Vision Sense
ICASSP 2025accepted
It is a new trend to fine-tune Large Multimodal Models (LMMs) to adapt to specific visual tasks through task-related conversation data. This approach provides a new paradigm for solving various vision-language tasks, however, it still faces two problems: (1) the global visual features input to the b…