← Search

Zailong Chen

2 accepted papers

2026

Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience Strategies

AAAI 2026technical

Large Vision-Language Models (LVLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they still face critical challenges in modeling long-range dependencies under the usage of Rotary Positional Encoding (ROPE). Although it can facilitate precise modeling of tok

Cited by 0SourcePDFScholar
2025

NCL-CIR: Noise-aware Contrastive Learning for Composed Image Retrieval

ICASSP 2025accepted

Composed Image Retrieval (CIR) seeks to find a target image using a multi-modal query, which combines an image with modification text to pinpoint the target. While recent CIR methods have shown promise, they mainly focus on exploring relationships between the query pairs (image and text) through dat…

Cited by 0SourceScholar