← Search

Zhongyu Yang

6 accepted papers

2026

Do Vision and Text Cues Exhibit Evidential Coupling? UFO: A Benchmark for Compositional Multimodal Reasoning in Unified Models

ICML 2026poster

Unified Foundation Models (UFMs), which support interleaved multimodal generation and understanding, have been proposed as a promising paradigm for reasoning about dynamic world states, yet it remains unclear whether the visual content they generate functions as grounded evidence for subsequent reas…

Cited by 0SourceScholar
2026

SVAgent: Storyline-guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration

CVPR 2026

Video question answering (VideoQA) is a challenging task that requires integrating spatial, temporal, and semantic information to capture the complex dynamics of video sequences. Although recent advances have introduced various approaches for video understanding, most existing methods still rely on

Cited by 0SourceScholar
2025

MERMAID: Multi-perspective Self-reflective Agents with Generative Augmentation for Emotion Recognition

EMNLP 2025

Multimodal large language models (MLLMs) have demonstrated strong performance across diverse multimodal tasks, achieving promising outcomes. However, their application to emotion recognition in natural images remains underexplored. MLLMs struggle to handle ambiguous emotional expressions and implici

Cited by 0SourcePDFScholar
2025

WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation

ICCV 2025poster

Knowledge discovery and collection are intelligence-intensive tasks that traditionally require significant human effort to ensure high-quality outputs. Recent research has explored multi-agent frameworks for automating Wikipedia-style article generation by retrieving and synthesizing information fro…

2023

CANet: Curved Guide Line Network with Adaptive Decoder for Lane Detection

ICASSP 2023accepted

Lane detection is challenging due to the complicated onroad scenarios and line deformation from different camera perspectives. Lots of solutions were proposed, but can not deal with "corner lanes" well. To address this problem, this paper proposes a new top-down deep learning lane detection approach…

Cited by 0SourceScholar