← Search

Yin Wu

4 accepted papers

2026

Causality Matters: How Temporal Information Emerges in Video Language Models

AAAI 2026technical

Video language models (VideoLMs) have made significant progress in multimodal understanding. However, temporal understanding, which involves identifying event order, duration, and relationships across time, still remains a core challenge. Prior works emphasize positional encodings (PEs) as a key mec

Cited by 0SourcePDFScholar
2024

MAR: Matching-Augmented Reasoning for Enhancing Visual-based Entity Question Answering

EMNLP 2024main

A multimodal large language model MLLMs may struggle with answering visual-based (personal) entity questions (VEQA), such as ”who is A?” or ”who is A that B is talking to?” for various reasons, e.g., the absence of the name of A in the caption or the inability of MLLMs to recognize A, particularly f…

2021

What the Role is vs. What Plays the Role: Semi-Supervised Event Argument Extraction via Dual Question Answering

AAAI 2021technical

Event argument extraction is an essential task in event extraction, and become particularly challenging in the case of low-resource scenarios. We solve the issues in existing studies under low-resource situations from two sides. From the perspective of the model, the existing methods always suffer f…