← Search

Jun Xue

10 accepted papers

2026

Geometry-based Schrödinger Bridges for Trustworthy Multimodal Fusion

ICML 2026poster

Real-world multimodal systems must be robust against low-quality data, such as sensor noise, incomplete multimodal data and conflicting inputs. However, existing trustworthy fusion methods rely on the model's own prediction confidence to judge data quality. This creates a circular dependency: when a…

Cited by 0SourceScholar
2026

Inconsistency-aware Multimodal Schrodinger Bridge for Deepfake Localization

CVPR 2026

Audio-visual deepfake localization demands interval-level outputs that serve as temporal evidence. Despite recent progress, symmetric fusion under single-sided or asynchronous forgeries propagates cross-modal noise, degrading high-precision localization. We present IaMSB, an inconsistency-aware mult

Cited by 0SourceScholar
2026

Multimodal Fusion via Self-Consistent Task-Gradient Fields

ICML 2026poster

Multimodal learning aims to preserve as much task-related information as possible from different inputs. However, current fusion designs often distort the feedback loop to feature extractors. Aggressively merging modalities entangles their representations, making the feature extractors fragile to in…

Cited by 0SourceScholar
2026

Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection

IJCAI 2026

The rapid advancement of generative AI has made audio deepfakes increasingly indistinguishable from authentic human vocals, posing significant threats to persons-of-interest (POI) such as public figures. Current detection systems primarily rely on generic, black-box models that fail to capture speak

Cited by 0Scholar
2026

When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse

CVPR 2026

Audio-Visual Speech Recognition (AVSR) has achieved remarkable progress in offline conditions, yet its robustness in real-world video conferencing (VC) remains largely unexplored. This paper presents the first systematic evaluation of state-of-the-art AVSR models across mainstream VC platforms, reve

Cited by 0SourceScholar
2025

MpoxMamba: A Grouped Mamba-based Lightweight Hybrid Network for Mpox Detection

ICASSP 2025accepted

Due to the lack of effective mpox detection tools, the mpox virus continues to spread worldwide and has been once again declared a public health emergency of international concern by the World Health Organization. Lightweight deep learning model-based detection systems are crucial to alleviate mpox…

Cited by 0SourceScholar
2025

Region-Based Optimization in Continual Learning for Audio Deepfake Detection

AAAI 2025technical

Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes when confronted with the diverse and evolving nature of real…

2024

Progressive Distillation Based on Masked Generation Feature Method for Knowledge Graph Completion

AAAI 2024technical

In recent years, knowledge graph completion (KGC) models based on pre-trained language model (PLM) have shown promising results. However, the large number of parameters and high computational cost of PLM models pose challenges for their application in downstream tasks. This paper proposes a progress…

2024

PseKD: Phase-Shift Encoded Knowledge Distillation for Oriented Object Detection in Remote Sensing Images

ICASSP 2024accepted

With the vigorous development of computer vision, oriented object detection has gradually been featured. However, angle boundary discontinuity and its knowledge distillation have been the bottleneck for rotating detection distillation design. In this paper, a novel knowledge distillation method name…

Cited by 0SourceScholar
2023

Learning From Yourself: A Self-Distillation Method For Fake Speech Detection

ICASSP 2023accepted

In this paper, we propose a novel self-distillation method for fake speech detection (FSD), which can significantly improve the performance of FSD without increasing the model complexity. For FSD, some fine-grained information is very important, such as spectrogram defects, mute segments, and so on,…

Cited by 0SourceScholar