← Search

Xiaoguang Zhu

6 accepted papers

2026

ROBUST MULTIMODAL REPRESENTATION LEARNING IN HEALTHCARE

ICASSP 2026poster

Medical multimodal representation learning aims to integrate heterogeneous data into unified patient representations to support clinical outcome prediction. However, real-world medical datasets commonly contain systematic biases from multiple sources, which poses significant challenges for medical m…

Cited by 0SourcePDFScholar
2026

ROBUST MULTIMODAL REPRESENTATION LEARNING IN HEALTHCARE

ICASSP 2026poster

Medical multimodal representation learning aims to integrate heterogeneous data into unified patient representations to support clinical outcome prediction. However, real-world medical datasets commonly contain systematic biases from multiple sources, which poses significant challenges for medical m…

Cited by 0SourcePDFScholar
2024

Supplementing Missing Visions Via Dialog for Scene Graph Generations

ICASSP 2024accepted

Most AI systems rely on the premise that the input visual data are sufficient to achieve competitive performance in various tasks. However, the classic task setup rarely considers the challenging, yet common practical situations where the complete visual data may be inaccessible due to various reaso…

Cited by 0SourceScholar
2023

Learning 3D Human Pose and Shape Estimation Using Uncertainty-Aware Body Part Segmentation

ICASSP 2023accepted

While exploiting body segmentations for supervision, existing 3D human pose and shape estimation methods are plagued by mismatches between clothed body segmentations and skinned SMPL model reprojections. Moreover, noisy pixels introduced by inaccurate segmentation annotations also prevent the model…

Cited by 0SourceScholar
2022

Learning Task-Specific Representation for Video Anomaly Detection with Spatial-Temporal Attention

ICASSP 2022accepted

The automatic detection of abnormal events in surveillance videos with weak supervision has been formulated as a multiple instance learning task, which aims to localize the clips containing abnormal events temporally with the video-level labels. However, most existing methods rely on the features ex…

Cited by 0SourceScholar
2019

Action Recognition Based on 3D Skeleton and RGB Frame Fusion

IROS 2019poster

Action recognition has wide applications in assisted living, health monitoring, surveillance, and human-computer interaction. In traditional action recognition methods, RGB video-based ones are effective but computationally inefficient, while skeleton-based ones are computationally efficient but do…

Cited by 40SourceScholar