← Search

Jiayu Xiong

4 accepted papers

2026

Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation

AAAI 2026technical

Multimodal Emotion Recognition in Conversation (MERC) aims to predict speakers’ emotions by integrating textual, acoustic, and visual cues. Existing approaches either struggle to capture complex cross‑modal interactions or experience gradient conflicts and unstable training when using deeper archite

Cited by 0SourcePDFScholar
2026

Geometry-based Schrödinger Bridges for Trustworthy Multimodal Fusion

ICML 2026poster

Real-world multimodal systems must be robust against low-quality data, such as sensor noise, incomplete multimodal data and conflicting inputs. However, existing trustworthy fusion methods rely on the model's own prediction confidence to judge data quality. This creates a circular dependency: when a…

Cited by 0SourceScholar
2026

Inconsistency-aware Multimodal Schrodinger Bridge for Deepfake Localization

CVPR 2026

Audio-visual deepfake localization demands interval-level outputs that serve as temporal evidence. Despite recent progress, symmetric fusion under single-sided or asynchronous forgeries propagates cross-modal noise, degrading high-precision localization. We present IaMSB, an inconsistency-aware mult

Cited by 0SourceScholar
2026

Multimodal Fusion via Self-Consistent Task-Gradient Fields

ICML 2026poster

Multimodal learning aims to preserve as much task-related information as possible from different inputs. However, current fusion designs often distort the feedback loop to feature extractors. Aggressively merging modalities entangles their representations, making the feature extractors fragile to in…

Cited by 0SourceScholar