ICASSP 2025accepted0 citations

EventLens: Enhancing Visual Commonsense Reasoning by Leveraging Event-Aware Pretraining and Cross-modal Linking

Mingjie Ma, Zhihuan Yu, Yichao Ma, Guohui Li, Zhong Yang

Abstract

Visual Commonsense Reasoning (VCR) is a cognitive task, challenging models to answer visual questions, and to explain the rationale behind their answers. While Large Language Models (LLMs) offer potential for this task, VCR’s complex scenes require specialized approaches to activate their commonsense reasoning abilities, as existing Multimodal LLMs struggle with VCR’s visual events and unique reference tags. To address these challenges, we propose EventLens, which enhances VCR through Event-Aware Pretraining and Cross-modal Linking in Supervised Fine-tuning. First, we introduce a new pretraining stage that emulates human cognitive processes to improve LLM comprehension of complex scenarios. Second, during supervised fine-tuning, we leverage reference tags to explicitly bridge RoI features with text, maintaining semantic integrity across modalities. Additionally, instruct prompts and task-specific adapters help integrate LLMs’ knowledge with new commonsense reasoning. Experimental results demonstrate competitive performance with state-of-the-art methods and ablation studies verify the effectiveness of proposed EventLens.

BibTeX
@inproceedings{icassp2025_eventlensenhanci,
  title = {EventLens: Enhancing Visual Commonsense Reasoning by Leveraging Event-Aware Pretraining and Cross-modal Linking},
  author = {Mingjie Ma and Zhihuan Yu and Yichao Ma and Guohui Li and Zhong Yang},
  booktitle = {ICASSP 2025},
  year = {2025}
}