← Search

Ze Zhou

2 accepted papers

2025

CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering

CVPR 2025highlight

Multimodal large language models (MLLMs) have garnered widespread attention from researchers due to their remarkable understanding and generation capabilities in visual language tasks (e.g., visual question answering). However, the rapid pace of knowledge updates in the real world makes offline trai…

Cited by 2SourcePDFScholar
2025

Relation-aware Semantic Alignment Network for Text-to-Image Person Retrieval

ICASSP 2025accepted

Text-to-Image Person Retrieval (TIPR) aims to utilize natural language descriptions as queries to retrieve pedestrian images. However, existing methods only concentrated on aligning individual text-image pairs and ignored the specific self-representations within both visible images and textual descr…

Cited by 0SourceScholar