← Search

Haoran Lou

3 accepted papers

2026

SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs

ICML 2026poster

Multimodal Large Language Models (MLLMs) possess intrinsic reasoning and world-knowledge capabilities, yet adapting them for dense retrieval remains challenging. Existing approaches typically rely on invasive parameter updates, such as full fine-tuning and LoRA, which risk disrupting the pre-trained…

Cited by 0SourceScholar
2025

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs

ICCV 2025poster

The architecture of multimodal large language models (MLLMs) commonly connects a vision encoder, often based on CLIP-ViT, to a large language model. While CLIP-ViT works well for capturing global image features, it struggles to model local relationships between adjacent patches, leading to weaker vi…

2025

MIND: A Multi-agent Framework for Zero-shot Harmful Meme Detection

ACL 2025long

The rapid expansion of memes on social media has highlighted the urgent need for effective approaches to detect harmful content. However, traditional data-driven approaches struggle to detect new memes due to their evolving nature and the lack of up-to-date annotated data. To address this issue, we…