← Search

Yuchen Xian

3 accepted papers

2026

From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion

ICML 2026poster

Multimodal image fusion (MMIF) aims to integrate complementary information from different modalities into a single fused image that preserves *fine local details* while maintaining *globally consistent appearance*. Most existing approaches build shared representations on 2D feature grids, which exce…

Cited by 0SourceScholar
2024

VISTA-LLAMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens

CVPR 2024poster

Recent advances in large video-language models have displayed promising outcomes in video comprehension. Current approaches straightforwardly convert video into language tokens and employ large language models for multi-modal tasks. However this method often leads to the generation of irrelevant con…

Cited by 16SourcePDFScholar