← Search

Mouxing Yang

16 accepted papers

2026

Endowing Vision-Language Models with System 2 Thinking for Fine-grained Visual Recognition

AAAI 2026technical

Vision-Language Models (VLMs) excel at extracting salient visual features from query images, thus exhibiting promising visual recognition performance. However, VLMs would encounter significant degradation in fine-grained scenarios due to their deficiency in distinguishing nuanced differences among c

Cited by 0SourcePDFScholar
2026

Learning Locally, Revising Globally: Global Reviser for Federated Learning with Noisy Labels

ICML 2026poster

In pursuit of data privacy, federated learning (FL) collaboratively trains a global model by aggregating local models learned from decentralized data. However, FL heavily depends on high-quality labels, which are often impractical in the real world, leading to the federated label-noise (F-LN) proble…

Cited by 0SourceScholar
2026

Learning with Dual-level Noisy Correspondence for Multi-modal Entity Alignment

ICLR 2026oral

Multi-modal entity alignment (MMEA) aims to identify equivalent entities across heterogeneous multi-modal knowledge graphs (MMKGs), where each entity is described by attributes from various modalities. Existing methods typically assume that both intra-entity and inter-graph correspondences are fault…

Cited by 0SourcecodeScholar
2026

Uncover Underlying Correspondence for Robust Multi-view Clustering

ICLR 2026oral

Multi-view clustering (MVC) aims to group unlabeled data into semantically meaningful clusters by leveraging cross-view consistency. However, real-world datasets collected from the web often suffer from noisy correspondence (NC), which breaks the consistency prior and results in unreliable alignmen…

Cited by 0SourcecodeScholar
2025

LLaVA-ReID: Selective Multi-image Questioner for Interactive Person Re-Identification

ICML 2025poster

Traditional text-based person ReID assumes that person descriptions from witnesses are complete and provided at once. However, in real-world scenarios, such descriptions are often partial or vague. To address this limitation, we introduce a new task called interactive person re-identification (Inter…

2025

Probabilistic Multimodal Learning with von Mises-Fisher Distributions

IJCAI 2025

Multimodal learning is pivotal for the advancement of artificial intelligence, enabling machines to integrate complementary information from diverse data sources for holistic perception and understanding. Despite significant progress, existing methods struggle with challenges such as noisy inputs, n

2025

Test-time Adaptation for Cross-modal Retrieval with Query Shift

ICLR 2025spotlight

The success of most existing cross-modal retrieval methods heavily relies on the assumption that the given queries follow the same distribution of the source domain. However, such an assumption is easily violated in real-world scenarios due to the complexity and diversity of queries, thus leading t…

2025

Visual Abstraction: A Plug-and-Play Approach for Text-Visual Retrieval

ICML 2025poster

Text-to-visual retrieval often struggles with semantic redundancy and granularity mismatches between textual queries and visual content. Unlike existing methods that address these challenges during training, we propose VISual Abstraction (VISA), a test-time approach that enhances retrieval by transf…

Cited by 0SourcePDFScholar
2024

Decoupled Contrastive Multi-View Clustering with High-Order Random Walks

AAAI 2024technical

In recent, some robust contrastive multi-view clustering (MvC) methods have been proposed, which construct data pairs from neighborhoods to alleviate the false negative issue, i.e., some intra-cluster samples are wrongly treated as negative pairs. Although promising performance has been achieved by…

2024

Robust Contrastive Multi-view Clustering against Dual Noisy Correspondence

NeurIPS 2024poster

Recently, contrastive multi-view clustering (MvC) has emerged as a promising avenue for analyzing data from heterogeneous sources, typically leveraging the off-the-shelf instances as positives and randomly sampled ones as negatives. In practice, however, this paradigm would unavoidably suffer from t…

2024

Test-time Adaptation against Multi-modal Reliability Bias

ICLR 2024poster

Test-time adaptation (TTA) has emerged as a new paradigm for reconciling distribution shifts across domains without accessing source data. However, existing TTA methods mainly concentrate on uni-modal tasks, overlooking the complexity of multi-modal scenarios. In this paper, we delve into the multi-…

2023

Graph Matching with Bi-level Noisy Correspondence

ICCV 2023poster

In this paper, we study a novel and widely existing problem in graph matching (GM), namely, Bi-level Noisy Correspondence (BNC), which refers to node-level noisy correspondence (NNC) and edge-level noisy correspondence (ENC). In brief, on the one hand, due to the poor recognizability and viewpoint d…

Cited by 42PDFcodeScholar
2023

Incomplete Multi-view Clustering via Prototype-based Imputation

IJCAI 2023poster

In this paper, we study how to achieve two characteristics highly-expected by incomplete multi-view clustering (IMvC). Namely, i) instance commonality refers to that within-cluster instances should share a common pattern, and ii) view versatility refers to that cross-view samples should own view-spe…

2022

Learning With Twin Noisy Labels for Visible-Infrared Person Re-Identification

CVPR 2022poster

In this paper, we study an untouched problem in visible-infrared person re-identification (VI-ReID), namely, Twin Noise Labels (TNL) which refers to as noisy annotation and correspondence. In brief, on the one hand, it is inevitable to annotate some persons with the wrong identity due to the complex…

Cited by 213PDFcodeScholar
2021

Partially View-Aligned Representation Learning With Noise-Robust Contrastive Loss

CVPR 2021poster

In real-world applications, it is common that only a portion of data is aligned across views due to spatial, temporal, or spatiotemporal asynchronism, thus leading to the so-called Partially View-aligned Problem (PVP). To solve such a less-touched problem without the help of labels, we propose simul…

Cited by 170PDFScholar