← Search

Shuai Zhong

5 accepted papers

2025

BIG-FUSION: Brain-Inspired Global-Local Context Fusion Framework for Multimodal Emotion Recognition in Conversations

AAAI 2025technical

Considering the importance of capturing both global conversational topics and local speaker dependencies for multimodal emotion recognition in conversations, current approaches first utilize sequence models like Transformer to extract global context information, then apply Graph Neural Networks to m…

Cited by 0SourcePDFScholar
2025

ClingTP: Curriculum Learning based Multi-style Title Prefix Generation

ICASSP 2025accepted

An informative, creative title prefix is memorable, capable of capturing the attention of readers, and significantly enhances the potential for increased citations. In this work, we pioneer the exploration of the significance of title prefixes in academic papers and propose a controllable title pref…

Cited by 0SourceScholar
2025

ImgTrojan: Jailbreaking Vision-Language Models with ONE Image

NAACL 2025long

There has been an increasing interest in the alignment of large language models (LLMs) with human values. However, the safety issues of their integration with a vision module, or vision language models (VLMs), remain relatively underexplored. In this paper, we propose a novel jailbreaking attack aga…

2025

Multi-View Incremental Learning with Structured Hebbian Plasticity for Enhanced Fusion Efficiency

AAAI 2025technical

The rapid evolution of multimedia technology has revolutionized human perception, paving the way for multi-view learning. However, traditional multi-view learning approaches are tailored for scenarios with fixed data views, falling short of emulating the intricate cognitive procedures of the human b…

Cited by 1SourcePDFScholar
2024

Exploring Question Guidance and Answer Calibration for Visually Grounded Video Question Answering

EMNLP 2024finding

Video Question Answering (VideoQA) tasks require not only correct answers but also visual evidence. The “localize-then-answer” strategy, while enhancing accuracy and interpretability, faces challenges due to the lack of temporal localization labels in VideoQA datasets. Existing methods often train t…

Cited by 0SourcePDFScholar