← Search

Mingjie Ma

2 accepted papers

2026

Seeing What Matters: A Training-Free Self-Guided Framework for Multimodal Detail Perception and Reasoning

CVPR 2026

Multimodal large language models (MLLMs) have achieved remarkable success on diverse visual-language tasks. However, fixed-resolution models face challenges in perceiving fine-grained visual details, particularly due to *distracted attention* and *blurry vision*. To address these issues, we propose

Cited by 0SourceScholar
2025

EventLens: Enhancing Visual Commonsense Reasoning by Leveraging Event-Aware Pretraining and Cross-modal Linking

ICASSP 2025accepted

Visual Commonsense Reasoning (VCR) is a cognitive task, challenging models to answer visual questions, and to explain the rationale behind their answers. While Large Language Models (LLMs) offer potential for this task, VCR’s complex scenes require specialized approaches to activate their commonsens…

Cited by 0SourceScholar