← Search

Zhihuan Yu

3 accepted papers

2025

EventLens: Enhancing Visual Commonsense Reasoning by Leveraging Event-Aware Pretraining and Cross-modal Linking

ICASSP 2025accepted

Visual Commonsense Reasoning (VCR) is a cognitive task, challenging models to answer visual questions, and to explain the rationale behind their answers. While Large Language Models (LLMs) offer potential for this task, VCR’s complex scenes require specialized approaches to activate their commonsens…

Cited by 0SourceScholar
2024

LMD: Faster Image Reconstruction with Latent Masking Diffusion

AAAI 2024technical

As a class of fruitful approaches, diffusion probabilistic models (DPMs) have shown excellent advantages in high-resolution image reconstruction. On the other hand, masked autoencoders (MAEs), as popular self-supervised vision learners, have demonstrated simpler and more effective image reconstructi…

2023

HybridPrompt: Bridging Language Models and Human Priors in Prompt Tuning for Visual Question Answering

AAAI 2023technical

Visual Question Answering (VQA) aims to answer the natural language question about a given image by understanding multimodal content. However, the answer quality of most existing visual-language pre-training (VLP) methods is still limited, mainly due to: (1) Incompatibility. Upstream pre-training ta…