← Search

Shaozu Yuan

12 accepted papers

2026

EmoThinker: Advancing Visual-Acoustic Emotion Analysis via Structural Token Selection and Chain-of-Thought Reasoning

CVPR 2026

Multimodal Emotion Analysis (MEA) is crucial for human-centric AI, yet current methods struggle with two core challenges: the sparse nature of emotional cues across modalities and their inherent temporal asynchrony. Existing approaches, which often rely on implicit fusion, consequently suffer from d

Cited by 0SourceScholar
2026

Tackling Model Bias via Game-theoretic Multi-agent Collaboration Framework for Hateful Meme Classification

CVPR 2026

Hateful meme classification aims to identify memes containing hateful content and has become increasingly important in the era of social media dominance. Large multimodal models (LMMs) have significantly enhanced the understanding of multimodal content, advancing this field. However, cognitive biase

Cited by 0SourcecodeScholar
2025

AutoMV: An Autonomous Agent Framework for Real Estate Marketing Video Generation

AAAI 2025technical

In this paper, we introduce AutoMV, an autonomous agent framework designed for generating real estate marketing videos. The framework integrates a diverse set of existing models into a tool library, allowing the agent to intelligently select and execute the appropriate tools. Given property images a…

Cited by 0SourcePDFScholar
2025

Multiple Feature Refining Network for Visual Emotion Distribution Learning

AAAI 2025technical

The significance of visual emotion distribution learning (VEDL) has surged, particularly with the growing inclination to convey emotions through images. The key of VEDL lies in capturing both low- and high-level features within the same visual content, thus promoting the model for salient and subtle…

2025

Towards Multimodal Sentiment Analysis via Hierarchical Correlation Modeling with Semantic Distribution Constraints

AAAI 2025technical

Sentiment analysis is rapidly advancing by utilizing various data modalities (e.g., text, video, and audio). However, most existing techniques only learn the atomic-level features that reflect strong correlations, while ignoring more complex compositions in multimodal data. Moreover, they also negle…

2024

G^2SAM: Graph-Based Global Semantic Awareness Method for Multimodal Sarcasm Detection

AAAI 2024technical

Multimodal sarcasm detection, aiming to detect the ironic sentiment within multimodal social data, has gained substantial popularity in both the natural language processing and computer vision communities. Recently, graph-based studies by drawing sentimental relations to detect multimodal sarcasm ha…

2023

Auslan-Daily: Australian Sign Language Translation for Daily Communication and News

NeurIPS 2023poster

Sign language translation (SLT) aims to convert a continuous sign language video clip into a spoken language. Considering different geographic regions generally have their own native sign languages, it is valuable to establish corresponding SLT datasets to support related communication and research.…

Cited by 18SourcePDFScholar
2023

Enhancing Multimodal Alignment with Momentum Augmentation for Dense Video Captioning

ICASSP 2023accepted

Dense video captioning aims to localize multiple events from an untrimmed video and generate corresponding captions for each event. Fusing different modalities(e.g. rgb, flow, audio) via transformer structure is a promising way to improve the caption performance. However, it is challenging for the c…

Cited by 0SourceScholar
2023

Nested Attention Network with Graph Filtering for Visual Question and Answering

ICASSP 2023accepted

Recently, Visual Question Answering(VQA), which is required to generate the answer by understanding both visual and textual content, has attracted considerable research interest. Most existing works extract visual features with the CNN network and learn its feature embedding with an attention mechan…

Cited by 0SourceScholar
2023

Tackling Modality Heterogeneity with Multi-View Calibration Network for Multimodal Sentiment Detection

ACL 2023long

With the popularity of social media, detecting sentiment from multimodal posts (e.g. image-text pairs) has attracted substantial attention recently. Existing works mainly focus on fusing different features but ignore the challenge of modality heterogeneity. Specifically, different modalities with in…

2022

Few-Shot Table Understanding: A Benchmark Dataset and Pre-Training Baseline

COLING 2022main

Few-shot table understanding is a critical and challenging problem in real-world scenario as annotations over large amount of tables are usually costly. Pre-trained language models (PLMs), which have recently flourished on tabular data, have demonstrated their effectiveness for table understanding t…

2022

Learning to Generate Poetic Chinese Landscape Painting with Calligraphy

IJCAI 2022poster

In this paper, we present a novel system (denoted as Polaca) to generate poetic Chinese landscape painting with calligraphy. Unlike previous single image-to-image painting generation, Polaca takes the classic poetry as input and outputs the artistic landscape painting image with the corresponding ca…

Cited by 10SourcePDFScholar