← Search

Hsu Tzu-Yin

1 accepted papers

2025

QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models

EMNLP 2025

Recently, Multimodal Large Language Models (MLLMs) encounter two key issues in multi-image contexts: (1) a lack of fine-grained perception across disparate images, and (2) a diminished capability to effectively reason over and synthesize information from multiple visual inputs. However, while variou

Cited by 0SourcePDFScholar