← Search

Xiaomin Song

8 accepted papers

2026

Multimodal Nested Learning for Decoupled and Coordinated Optimization

ICML 2026oral

Multimodal learning aims to integrate multi-sensor data to exploit their complementary information, embracing a more comprehensive real-world perception and understanding. However, heterogeneous discrepancies across modalities consistently trigger imbalanced multimodal optimization, restricting the …

Cited by 0SourceScholar
2026

Robust Semi-paired Multimodal Learning for Cross-modal Retrieval

AAAI 2026technical

Cross-modal retrieval is a fundamental application of multi-modal learning that has achieved remarkable success with large-scale well-paired data. However, in practice, it is costly to collect large-scale well-paired data. To alleviate the dependence on the amount of paired data, in this paper, we s

Cited by 0SourcePDFScholar
2025

Fuzzy Multimodal Learning for Trusted Cross-modal Retrieval

CVPR 2025poster

Cross-modal retrieval aims to match related samples across distinct modalities, facilitating the retrieval and discovery of heterogeneous information. Although existing methods show promising performance, most are deterministic models and are unable to capture the uncertainty inherent in the retriev…

2025

RoDA: Robust Domain Alignment for Cross-Domain Retrieval Against Label Noise

AAAI 2025technical

This paper studies the complex challenge of cross-domain image retrieval under the condition of noisy labels (NCIR), a scenario that not only includes the inherent obstacles of traditional cross-domain image retrieval (CIR) but also requires alleviating the adverse effects of label noise. To address…

2025

Robust Cross-modal Alignment Learning for Cross-Scene Spatial Reasoning and Grounding

NeurIPS 2025poster

Grounding target objects in 3D environments via natural language is a fundamental capability for autonomous agents to successfully fulfill user requests. Almost all existing works typically assume that the target object lies within a known scene and focus solely on in-scene localization. In practice…

Cited by 0SourceScholar
2025

Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy Labels

AAAI 2025technical

Cross-modal hashing (CMH) has appeared as a popular technique for cross-modal retrieval due to its low storage cost and high computational efficiency in large-scale data. Most existing methods implicitly assume that multi-modal data is correctly labeled, which is expensive and even unattainable due…

2021

Time Series Data Augmentation for Deep Learning: A Survey

IJCAI 2021poster

Deep learning performs remarkably well on many time series analysis tasks recently. The superior performance of deep neural networks relies heavily on a large number of training data to avoid overfitting. However, the labeled data of many real-world time series applications may be limited such as cl…