← Search

Guiduo Duan

11 accepted papers

2026

LangPrecip: Language-Aware Multimodal Precipitation Nowcasting

ICML 2026poster

Short-term precipitation nowcasting is inherently under-constrained due to limited historical observation windows: identical observations can lead to multiple plausible future trajectories, especially for extreme events. Existing generative methods rely solely on visual features and lack explicit co…

Cited by 0SourceScholar
2026

TiCAL:Typicality-Based Consistency-Aware Learning for Multimodal Emotion Recognition

AAAI 2026technical

Multimodal Emotion Recognition (MER) aims to accurately identify human emotional states by integrating heterogeneous modalities such as visual, auditory, and textual data. Existing approaches predominantly rely on unified emotion labels to supervise model training, often overlooking a critical chall

Cited by 0SourcePDFScholar
2026

Uncertainty-Gated Deformable Network for Breast Tumor Segmentation in MR images

ICASSP 2026poster

Accurate segmentation of breast tumors in magnetic resonance images (MRI) is essential for breast cancer diagnosis, yet existing methods face challenges in capturing irregular tumor shapes and effectively integrating local and global features. To address these limitations, we propose an uncertainty-…

Cited by 0SourcePDFScholar
2025

DFMA: Adaptive Dual Fusion for Multimodal Relation Extraction with Mutual Attention

ICASSP 2025accepted

Multimodal relation extraction (MRE) is an emerging research field that combines techniques from natural language processing, computer vision, and machine learning, helping us better understand and interpret data. However, current methods are faced with two main issues. The first issue is that the a…

Cited by 0SourceScholar
2025

Deep Fuzzy Multi-view Learning for Reliable Classification

ICML 2025poster

Multi-view learning methods primarily focus on enhancing decision accuracy but often neglect the uncertainty arising from the intrinsic drawbacks of data, such as noise, conflicts, etc. To address this issue, several trusted multi-view learning approaches based on the Evidential Theory have been pro…

Cited by 0SourcePDFScholar
2025

Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition

CVPR 2025poster

Visual Emotion Recognition (VER) is a critical yet challenging task aimed at inferring emotional states of individuals based on visual cues. However, existing works focus on single domains, e.g., realistic images or stickers, limiting VER models' cross-domain generalizability. To fill this gap, we…

Cited by 0SourcePDFScholar
2025

ROLL: Robust Noisy Pseudo-label Learning for Multi-View Clustering with Noisy Correspondence

CVPR 2025highlight

Multi-view clustering (MVC) aims to exploit complementary information from diverse views to enhance clustering performance. Since pseudo-labels can provide additional semantic information, many MVC methods have been proposed to guide unsupervised multi-view learning through pseudo-labels. These meth…

Cited by 0SourcePDFScholar
2025

Reliable Disentanglement Multi-view Learning Against View Adversarial Attacks

IJCAI 2025

Trustworthy multi-view learning has attracted extensive attention because evidence learning can provide reliable uncertainty estimation to enhance the credibility of multi-view predictions. Existing trusted multi-view learning methods implicitly assume that multi-view data is secure. However, in saf

2025

SPADE: Spatial-Aware Denoising Network for Open-vocabulary Panoptic Scene Graph Generation with Long- and Local-range Context Reasoning

ICCV 2025poster

Panoptic Scene Graph Generation (PSG) integrates instance segmentation with relation understanding to capture pixel-level structural relationships in complex scenes. Although recent approaches leveraging pre-trained vision-language models (VLMs) have significantly improved performance in the open-vo…

Cited by 0SourcePDFScholar
2025

WMAJL: Watcher-Mediated Attention Joint Learning Model for Multimodal Relation Extraction

ICASSP 2025accepted

In the domain of Multimodal Relation Extraction (MRE), we present the $\color{Red}{\text{W}}$atcher-$\color{Red}{\text{M}}$ediated $\color{Red}{\text{A}}$ttention $\color{Red}{\text{J}}$oint $\color{Red}{\text{L}}$earning Model ($\color{Red}{\text{WMAJL}}$), a novel approach addressing the challenge…

Cited by 0SourceScholar