← Search

Siyang Song

30 accepted papers

2026

Explainable Depression Assessment from Face Videos by Weakly Supervised Learning

AAAI 2026technical

Existing video-based automatic depression assessment (ADA) approaches frequently achieve video-level depression assessment by aggregating features or predictions of individual frames or equal-length segments within the given video. While their performances have been largely enhanced by recent advanc

Cited by 0SourcePDFScholar
2026

MAUGen: A Unified Diffusion Approach for Multi-Identity Facial Expression and AU Label Generation

AAAI 2026technical

The lack of large-scale, demographically diverse face images with precise Action Unit (AU) occurrence and intensity annotations has long been recognized as a fundamental bottleneck in developing generalizable facial AU recognition systems. In this paper, we propose MAUGen, a diffusion-based multi-mo

Cited by 0SourcePDFScholar
2026

PureCC: Pure Learning for Text-to-Image Concept Customization

CVPR 2026

Existing concept customization methods have achieved remarkable outcomes in high-fidelity and multi-concept customization. However, they often neglect the influence on the original model's behavior and capabilities when learning new personalized concepts. To address this issue, we propose PureCC. Pu

Cited by 0SourcecodeScholar
2025

A Frequency-aware Augmentation Network for Mental Disorders Assessment from Audio

ICASSP 2025accepted

Depression and Attention Deficit Hyperactivity Disorder (ADHD) stand out as the common mental health challenges today. In affective computing, speech signals serve as effective biomarkers for mental disorder assessment. Current research, relying on labor-intensive hand-crafted features or simplistic…

Cited by 0SourceScholar
2025

CA-Edit: Causality-Aware Condition Adapter for High-Fidelity Local Facial Attribute Editing

AAAI 2025technical

For efficient and high-fidelity local facial attribute editing, most existing editing methods either require additional fine-tuning for different editing effects or tend to affect beyond the editing regions. Alternatively, inpainting methods can edit the target image region while preserving external…

2025

DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face Synthesis

ICASSP 2025accepted

Accurately synthesizing talking face videos and capturing fine facial features for individuals with long hair presents a significant challenge. To tackle these challenges in existing methods, we propose a decomposed per-embedding Gaussian fields (DEGSTalk), a 3D Gaussian Splatting (3DGS)-based talki…

Cited by 0SourceScholar
2025

DepMGNN: Matrixial Graph Neural Network for Video-based Automatic Depression Assessment

AAAI 2025technical

Depression can be reflected by long-term human spatio-temporal facial behaviours. While human face videos recorded in real-world usually have long and variable lengths, existing video-based depression assessment approaches frequently re-sample/down-sample such videos to short and equal-length videos…

2025

Hierarchical Multimodal Decoupling-Fusion Framework for offline Multiple Appropriate Facial Reaction Generation

ICASSP 2025accepted

Facial reactions convey crucial emotional information and coordinating interpersonal relationships in human dyadic interactions. While existing Multiple Appropriate Facial Reaction Generation (MAFRG) methods focus on generating multiple reasonable facial reactions, none of these approaches combines…

Cited by 0SourceScholar
2025

Learning from Human Conversations: A Seq2Seq based Multi-modal Robot Facial Expression Reaction Framework in HRI

IROS 2025

Nonverbal communication plays a crucial role in both human-human and human-robot interactions (HRIs), where facial expressions convey emotions, intentions and trust. Enabling humanoid robots to generate human-like facial reactions in response to human speech and facial behaviours remains significant

Cited by 0SourcecodeScholar
2025

M3ADD: A Novel Benchmark for Physiology Signal-based Automatic Depression Detection with Multimodal Multitask Multievent Framework

ICASSP 2025accepted

The prevalence of depression is escalating, especially among youth, which has become a critical mental health concern. Current assessment methods, relying heavily on questionnaires, clinical observations, and AI-driven analyses, are limited by their focus on single-event data, failing to encapsulate…

Cited by 0SourceScholar
2025

MERMAID: Multi-perspective Self-reflective Agents with Generative Augmentation for Emotion Recognition

EMNLP 2025

Multimodal large language models (MLLMs) have demonstrated strong performance across diverse multimodal tasks, achieving promising outcomes. However, their application to emotion recognition in natural images remains underexplored. MLLMs struggle to handle ambiguous emotional expressions and implici

Cited by 0SourcePDFScholar
2025

MSAmba: Exploring Multimodal Sentiment Analysis with State Space Models

AAAI 2025technical

Multimodal sentiment analysis, which learns a model to process multiple modalities simultaneously and predict a sentiment value, is an important area of affective computing. Modeling sequential intra-modal information and enhancing cross-modal interactions are crucial to multimodal sentiment analysi…

2025

OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions

NeurIPS 2025poster

In this paper, we introduce Online Multimodal Conversational Response Generation (OMCRG), a novel task designed to produce synchronized verbal and non-verbal listener feedback online, based on the speaker's multimodal inputs. OMCRG captures natural dyadic interactions and introduces new challenges i…

Cited by 0SourcecodeScholar
2025

PerReactor: Offline Personalised Multiple Appropriate Facial Reaction Generation

AAAI 2025technical

In dyadic human-human interactions, individuals may express multiple different facial reactions in response to the same/similar behaviours expressed by their conversational partners depending on their personalised behaviour patterns. As a result, frequently-employed reconstruction loss-based strateg…

2025

SynFER: Towards Boosting Facial Expression Recognition with Synthetic Data

ICCV 2025poster

Facial expression datasets remain limited in scale due to privacy concerns, the subjectivity of annotations, and the labor-intensive nature of data collection. This limitation poses a significant challenge for developing modern deep learning-based facial expression analysis models, particularly foun…

Cited by 0SourcePDFScholar
2024

Advancing Saliency Ranking with Human Fixations: Dataset Models and Benchmarks

CVPR 2024poster

Saliency ranking detection (SRD) has emerged as a challenging task in computer vision aiming not only to identify salient objects within images but also to rank them based on their degree of saliency. Existing SRD datasets have been created primarily using mouse-trajectory data which inadequately ca…

2024

Boosting Adversarial Transferability across Model Genus by Deformation-Constrained Warping

AAAI 2024technical

Adversarial examples generated by a surrogate model typically exhibit limited transferability to unknown target systems. To address this problem, many transferability enhancement approaches (e.g., input transformation and model augmentation) have been proposed. However, they show poor performances i…

2024

CemiFace: Center-based Semi-hard Synthetic Face Generation for Face Recognition

NeurIPS 2024poster

Privacy issue is a main concern in developing face recognition techniques. Although synthetic face images can partially mitigate potential legal risks while maintaining effective face recognition (FR) performance, FR models trained by face images synthesized by existing generative approaches frequen…

2024

Circular Decomposition and Cross-Modal Recombination for Multimodal Sentiment Analysis

ICASSP 2024accepted

Multimodal Sentiment Analysis is a burgeoning research area, leveraging various modalities to predict the sentiment score. Nevertheless, previous studies have disregarded the impact of noise interference on specific modal sentiments during video recording, thereby compromising the accuracy of sentim…

Cited by 0SourceScholar
2024

Domain Separation Graph Neural Networks for Saliency Object Ranking

CVPR 2024poster

Saliency object ranking (SOR) has attracted significant attention recently. Previous methods usually failed to explicitly explore the saliency degree-related relationships between objects. In this paper we propose a novel Domain Separation Graph Neural Network (DSGNN) which starts with separately ex…

2024

MERG: Multi-Dimensional Edge Representation Generation Layer for Graph Neural Networks

ICASSP 2024accepted

Edges are essential in describing relationships among nodes. While existing graphs frequently use a single-value edge to describe association between each pair of node vectors, crucial relationships may be disregarded if they are not linearly correlated, which may limit graph analysis performance. A…

Cited by 0SourceScholar
2024

MTaDCS: Moving Trace and Feature Density-based Confidence Sample Selection under Label Noise

ECCV 2024poster

"Learning from noisy labels is a challenging task, as noisy labels can compromise decision boundaries and result in suboptimal generalization performance. Most previous approaches for dealing noisy labels are based on sample selection, which utilized the small loss criterion to reduce the adverse ef…

2024

Multi-Level Graph Learning For Audio Event Classification And Human-Perceived Annoyance Rating Prediction

ICASSP 2024accepted

WHO’s report on environmental noise estimates that 22 M people suffer from chronic annoyance related to noise caused by audio events (AEs) from various sources. Annoyance may lead to health issues and adverse effects on metabolic and cognitive systems. In cities, monitoring noise levels does not pro…

Cited by 0SourceScholar
2024

Scale-Free And Task-Generic Attack: Generating Photo-Realistic Adversarial Patterns With Patch Quilting Generator

ICASSP 2024accepted

Recent CNN generator-based attack approaches can synthe-size unrestricted and semantically meaningful entities to the image, which are able to improve the transferability and robustness. However, such methods attack images by either synthesizing local adversarial entities, which are only suitable fo…

Cited by 0SourceScholar
2024

Towards Combating Frequency Simplicity-biased Learning for Domain Generalization

NeurIPS 2024poster

Domain generalization methods aim to learn transferable knowledge from source domains that can generalize well to unseen target domains. Recent studies show that neural networks frequently suffer from a simplicity-biased learning behavior which leads to over-reliance on specific frequency sets, nam…

2023

Fourier-Net: Fast Image Registration with Band-Limited Deformation

AAAI 2023technical

Unsupervised image registration commonly adopts U-Net style networks to predict dense displacement fields in the full-resolution spatial domain. For high-resolution volumetric image data, this process is however resource-intensive and time-consuming. To tackle this problem, we propose the Fourier-Ne…

2023

Shift from Texture-bias to Shape-bias: Edge Deformation-based Augmentation for Robust Object Recognition

ICCV 2023poster

Recent studies have shown the vulnerability of CNNs under perturbation noises, which is partially caused by the reason that the well-trained CNNs are too biased toward the object texture, i.e., they make predictions mainly based on texture cues. To reduce this texture-bias, current studies resort to…

Cited by 7PDFcodeScholar
2022

Learning Multi-dimensional Edge Feature-based AU Relation Graph for Facial Action Unit Recognition

IJCAI 2022poster

The activations of Facial Action Units (AUs) mutually influence one another. While the relationship between a pair of AUs can be complex and unique, existing approaches fail to specifically and explicitly represent such cues for each pair of AUs in each facial display. This paper proposes an AU rela…

2022

Statistical, Spectral and Graph Representations for Video-Based Facial Expression Recognition in Children

ICASSP 2022accepted

Child facial expression recognition is a relatively less investigated area within affective computing. Children’s facial expressions differ significantly from adults; thus, it is necessary to develop emotion recognition frameworks that are more objective, descriptive and specific to this target user…

Cited by 0SourceScholar