← Search

Xiao Ke

10 accepted papers

2026

LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal Observations

ICML 2026spotlight

Real-world multimodal learning is often hindered by missing modalities. While Incomplete Multimodal Learning (IML) has gained traction, existing methods typically rely on the unrealistic assumption of full-modal availability during training to provide reconstruction supervision or cross-modal priors…

Cited by 0SourceScholar
2026

LaRA-Fusion: Latent-Robust Adaptation via Dual-Loop Constraints for Infrared and Visible Image Fusion

ICML 2026poster

Infrared and visible image fusion (IVIF) aims to synergize complementary thermal radiation and textural details for comprehensive scene perception. However, existing unsupervised paradigms often overlook the intrinsic topological consistency shared across modalities. Lacking explicit geometric regul…

Cited by 0SourceScholar
2026

MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality Assessment

AAAI 2026technical

Multimodal Action Quality Assessment (AQA) has recently emerged as a promising paradigm. By leveraging complementary information across shared contextual cues, it enhances the discriminative evaluation of subtle intra-class variations in highly similar action sequences. However, partial modalities a

Cited by 0SourcePDFScholar
2025

DanceFix: An Exploration in Group Dance Neatness Assessment Through Fixing Abnormal Challenges of Human Pose

AAAI 2025technical

The fair and objective assessment of performances and competitions is a common pursuit and challenge in human society. The application of computer vision technology offers hope for this purpose, but it still faces obstacles such as occlusion and motion blur. To address these hindrances, our DanceFix…

Cited by 0SourcePDFScholar
2025

Language-Guided Audio-Visual Learning for Long-Term Sports Assessment

CVPR 2025poster

Long-term sports assessment is a challenging task in video understanding since it requires judging complex movement variations and action-music coordination. However, there is no direct correlation between the diverse background music and movements in sporting events. Previous works require a large…

2025

Progressive Modality-Adaptive Interactive Network for Multi-Modality Image Fusion

IJCAI 2025

Multi-modality image fusion (MMIF) integrates features from distinct modalities to enhance visual quality and improve downstream task performance. However, existing methods often overlook the sparsity variations and dynamic correlations between infrared and visible images, potentially limiting the u

Cited by 0SourcePDFScholar
2025

Projection, Interaction and Fusion: A Progressive Difference Fusion Network for Salient Object Detection

IJCAI 2025

In recent years, deep learning-based Salient Object Detection (SOD) methods have made tremendous progress; however, their performance in complex scenarios has reached a bottleneck. In this paper, we propose a novel Progressive Difference Fusion Network (PDFNet) based on fine-grained feature fusion.

2024

Cutransnet: Transformers to Make Strong Encoders for Multi-Task Vision Perception of Autonomous Driving

ICASSP 2024accepted

In autonomous driving, perception plays a critical role as it serves as a fundamental requirement for both planning and control. Currently, most perception tasks are processed independently, which requires designing multiple models and networks to handle multiple tasks. This division leads to multip…

Cited by 0SourceScholar
2024

IF-Font: Ideographic Description Sequence-Following Font Generation

NeurIPS 2024poster

Few-shot font generation (FFG) aims to learn the target style from a limited number of reference glyphs and generate the remaining glyphs in the target font. Previous works focus on disentangling the content and style features of glyphs, combining the content features of the source glyph with the st…