← Search

Zhe Xue

14 accepted papers

2026

HFR-MKGC: Hierarchical Fusion Reasoning with MLLMs for Multi-modal Knowledge Graph Completion

AAAI 2026technical

Multi-modal knowledge graph completion (MMKGC) aims to infer missing entities of triples by leveraging heterogeneous information in knowledge graph (KG). However, existing approaches often struggle with inconsistent modality alignment, limited reasoning depth, and insufficient negative sample qualit

Cited by 0SourcePDFScholar
2026

POGA: Paraphrased and Oppositional Graph Alignment for Fine-Grained Cross-Modal Retrieval

CVPR 2026

Most of the models used to generate embeddings for retrieval are not trained for the purpose which leads them to focus on coarse semantic alignment rather than particular object attributes or arrangements. This limits their performance, particularly on challenging problems such as cross-modal fine-g

Cited by 0SourceScholar
2026

Rethink Representation Learning for Questionnaire Data

AAAI 2026technical

Questionnaire data serve as a valuable resource across numerous scientific domains, offering insights into human behavior, health, and social trends. Traditional downsampling-based representation learning methods—such as standardization and one-hot encoding—reformat these data into tabular structure

Cited by 0SourcePDFScholar
2026

ST-VLM: A Spatial-to-Image Multimodal Spatial-Temporal Prediction Framework with Vision-Language Model

AAAI 2026technical

Spatial-temporal prediction plays a crucial role in various domains, including intelligent transportation and environmental monitoring. Although large language model has shown advantages in long-range dependency modeling and excellent generalization ability for forecasting, it has limited understand

Cited by 0SourcePDFScholar
2025

Generalizing Single-Frame Supervision to Event-Level Understanding for Video Anomaly Detection

NeurIPS 2025poster

Video Anomaly Detection (VAD) aims to identify abnormal frames from discrete events within video sequences. Existing VAD methods suffer from heavy annotation burdens in fully-supervised paradigm, insensitivity to subtle anomalies in semi-supervised paradigm, and vulnerability to noise in weakly-supe…

Cited by 0SourceScholar
2025

Generating Synthetic Data for Unsupervised Federated Learning of Cross-Modal Retrieval

AAAI 2025technical

Unsupervised federated learning for cross-modal retrieval has received increasing attention in recent years as it can free the requirement for annotations and avoid uploading original clients’ data to servers. Most existing methods focus on how to learn better local models and their aggregation to o…

Cited by 0SourcePDFScholar
2025

Incomplete Multi-View Multi-Label Classification via Diffusion-Guided Redundancy Removal

AAAI 2025technical

Incomplete multi-view multi-label classification aims to accurately predict labels for each sample in the face of some missing views. Due to its widespread presence in real-world scenarios, it has become an extensively researched topic. In addition to the challenges brought by missing views, it also…

Cited by 0SourcePDFScholar
2025

Medusa: A Multi-Scale High-order Contrastive Dual-Diffusion Approach for Multi-View Clustering

CVPR 2025poster

Deep multi-view clustering methods utilize information from multiple views to achieve enhanced clustering results and have gained increasing popularity in recent years. Most existing methods typically focus on either inter-view or intra-view relationships, aiming to align information across views or…

Cited by 0SourcePDFScholar
2025

Reinforcement Active Client Selection for Federated Heterogeneous Graph Learning

AAAI 2025technical

Carefully selecting clients to participate in aggregation can assist the global model in achieving better performance. However, existing research on federated heterogeneous graph learning (FHGL) has shown limited attention to the client selection (CS) problem. Current CS algorithms face challenges i…

Cited by 0SourcePDFScholar
2024

Efficient Asynchronous Federated Learning with Prospective Momentum Aggregation and Fine-Grained Correction

AAAI 2024technical

Asynchronous federated learning (AFL) is a distributed machine learning technique that allows multiple devices to collaboratively train deep learning models without sharing local data. However, AFL suffers from low efficiency due to poor client model training quality and slow server model convergenc…

Cited by 9SourcePDFScholar
2024

Self-Supervised Multi-Modal Knowledge Graph Contrastive Hashing for Cross-Modal Search

AAAI 2024technical

Deep cross-modal hashing technology provides an effective and efficient cross-modal unified representation learning solution for cross-modal search. However, the existing methods neglect the implicit fine-grained multimodal knowledge relations between these modalities such as when the image contains…

Cited by 7SourcePDFScholar
2024

View-Category Interactive Sharing Transformer for Incomplete Multi-View Multi-Label Learning

CVPR 2024highlight

As a problem often encountered in real-world scenarios multi-view multi-label learning has attracted considerable research attention. However due to oversights in data collection and uncertainties in manual annotation real-world data often suffer from incompleteness. Regrettably most existing multi-…

Cited by 6SourcePDFScholar
2023

All in a Row: Compressed Convolution Networks for Graphs

ICML 2023poster

Compared to Euclidean convolution, existing graph convolution methods generally fail to learn diverse convolution operators under limited parameter scales and depend on additional treatments of multi-scale feature extraction. The challenges of generalizing Euclidean convolution to graphs arise from…

2021

Clustering-Induced Adaptive Structure Enhancing Network for Incomplete Multi-View Data

IJCAI 2021poster

Incomplete multi-view clustering aims to cluster samples with missing views, which has drawn more and more research interest. Although several methods have been developed for incomplete multi-view clustering, they fail to extract and exploit the comprehensive global and local structure of multi-view…

Cited by 41SourcePDFScholar