← Search

Xudong Yan

9 accepted papers

2026

Bridging the Modality Gap in Compositional Zero-Shot Learning via Sparse Alignment and Unimodal Memory Bank

CVPR 2026

Compositional Zero-Shot Learning (CZSL) aims to recognize unseen attribute-object compositions with learned primitives (attribute and object) knowledge from seen compositions. While previous approaches gain their notable performance through the powerful cross-modal alignment of CLIP, they often over

Cited by 0SourceScholar
2025

Leveraging MLLM Embeddings and Attribute Smoothing for Compositional Zero-Shot Learning

IJCAI 2025

Compositional zero-shot learning (CZSL) aims to recognize novel compositions of attributes and objects learned from seen compositions. Previous works disentangle attributes and objects by extracting shared and exclusive parts between the image pair sharing the same attribute (object), as well as ali

2025

MaxSup: Overcoming Representation Collapse in Label Smoothing

NeurIPS 2025oral

Label Smoothing (LS) is widely adopted to reduce overconfidence in neural network predictions and improve generalization. Despite these benefits, recent studies reveal two critical issues with LS. First, LS induces overconfidence in misclassified samples. Second, it compacts feature representations…

Cited by 0SourcecodeScholar
2025

TOMCAT: Test-time Comprehensive Knowledge Accumulation for Compositional Zero-Shot Learning

NeurIPS 2025poster

Compositional Zero-Shot Learning (CZSL) aims to recognize novel attribute-object compositions based on the knowledge learned from seen ones. Existing methods suffer from performance degradation caused by the distribution shift of label space at test time, which stems from the inclusion of unseen co…

Cited by 0SourcecodeScholar
2024

BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition

CVPR 2024poster

Graph Convolutional Networks (GCNs) have long set the state-of-the-art in skeleton-based action recognition leveraging their ability to unravel the complex dynamics of human joint topology through the graph's adjacency matrix. However an inherent flaw has come to light in these cutting-edge models:…

2023

Social Relation Reasoning Based on Triangular Constraints

AAAI 2023technical

Social networks are essentially in a graph structure where persons act as nodes and the edges connecting nodes denote social relations. The prediction of social relations, therefore, relies on the context in graphs to model the higher-order constraints among relations, which has not been exploited s…

Cited by 9SourcePDFScholar
2023

Visual Traffic Knowledge Graph Generation from Scene Images

ICCV 2023poster

Although previous works on traffic scene understanding have achieved great success, most of them stop at a lowlevel perception stage, such as road segmentation and lane detection, and few concern high-level understanding. In this paper, we present Visual Traffic Knowledge Graph Generation (VTKGG), a…

Cited by 15PDFScholar
2021

Learning Semantic Context from Normal Samples for Unsupervised Anomaly Detection

AAAI 2021technical

Unsupervised anomaly detection aims to identify data samples that have low probability density from a set of input samples, and only the normal samples are provided for model training. The inference of abnormal regions on the input image requires an understanding of the surrounding semantic context.…

Cited by 179SourcePDFScholar