← Search

Mozhgan Nasr Azadani

2 accepted papers

2025

Hawaii: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language Models

NeurIPS 2025poster

Improving the visual understanding ability of vision-language models (VLMs) is crucial for enhancing their performance across various tasks. While using multiple pretrained visual experts has shown great promise, it often incurs significant computational costs during training and inference. To addre…

Cited by 0SourceScholar
2025

LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts

EMNLP 2025

Redundancy of visual tokens in multi-modal large language models (MLLMs) significantly reduces their computational efficiency. Recent approaches, such as resamplers and summarizers, have sought to reduce the number of visual tokens, but at the cost of visual reasoning ability. To address this, we pr

Cited by 0SourcePDFScholar