← Search

Xiaojun Wu

27 accepted papers

2026

BAG: Benchmarking Anomaly Detection on Dynamic Graphs

AAAI 2026technical

Anomaly detection in dynamic graphs is a critical area of research that focuses on identifying abnormal components within evolving graph structures that deviate significantly from typical patterns. Despite advancements in traditional temporal pattern mining and deep learning techniques, a comprehens

Cited by 0SourcePDFScholar
2026

Beyond Strict Pairing: Arbitrarily Paired Training for High-Performance Infrared and Visible Image Fusion

CVPR 2026

Infrared and visible image fusion (IVIF) aims to synthesise complementary information from the two source modalities while preserving natural textures and salient thermal signatures simultaneously. Existing solutions predominantly rely on extensive sets of rigidly aligned image pairs for training. H

Cited by 0SourcecodeScholar
2026

DecompGAIL: Learning Realistic Traffic Behaviors with Decomposed Multi-Agent Generative Adversarial Imitation Learning

ICLR 2026poster

Realistic traffic simulation is critical for the development of autonomous driving systems and urban mobility planning, yet existing imitation learning approaches often fail to model realistic traffic behaviors. Behavior cloning suffers from covariate shift, while Generative Adversarial Imitation Le…

Cited by 0SourceScholar
2026

Fast and Stable Riemannian Metrics on SPD Manifolds via Cholesky Product Geometry

ICLR 2026poster

Recent advances in Symmetric Positive Definite (SPD) matrix learning show that Riemannian metrics are fundamental to effective SPD neural networks. Motivated by this, we revisit the geometry of the Cholesky factors and uncover a simple product structure that enables convenient metric design. Buildin…

Cited by 0SourcecodeScholar
2026

Text-Driven Fusion for Infrared and Visible Images: Achieving Image Scene Adaptation on Hyperbolic Space

ICML 2026poster

Infrared and visible image fusion aims to integrate complementary information from both modalities. However, most existing methods rely on Euclidean representations, which inherently impose geometric constraints that hinder effective semantic modelling. Specifically, Euclidean geometry imposes rigid…

Cited by 0SourceScholar
2026

Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interaction

CVPR 2026

With the rapid advancement and widespread application of vision-language pre-training (VLP) models, their vulnerability to adversarial attacks has become a critical concern. In general, the adversarial examples can typically be designed to exhibit transferable power, attacking not only different mod

Cited by 0SourcecodeScholar
2025

Adaptive Hyper-Graph Convolution Network for Skeleton-based Human Action Recognition with Virtual Connections

ICCV 2025poster

The shared topology of human skeletons motivated the recent investigation of graph convolutional network (GCN) solutions for action recognition.However, most of the existing GCNs rely on the binary connection of two neighboring vertices (joints) formed by an edge (bone), overlooking the potential of…

2025

Collaborating Vision, Depth, and Thermal Signals for Multi-Modal Tracking: Dataset and Algorithm

NeurIPS 2025poster

Existing multi-modal object tracking approaches primarily focus on dual-modal paradigms, such as RGB-Depth or RGB-Thermal, yet remain challenged in complex scenarios due to limited input modalities. To address this gap, this work introduces a novel multi-modal tracking task that leverages three com…

Cited by 0SourcecodeScholar
2025

Golden Touchstone: A Comprehensive Bilingual Benchmark for Evaluating Financial Large Language Models

EMNLP 2025

As large language models (LLMs) increasingly permeate the financial sector, there is a pressing need for a standardized method to comprehensively assess their performance. Existing financial benchmarks often suffer from limited language and task coverage, low-quality datasets, and inadequate adaptab

2025

Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated Learning

ICML 2025poster

Batch Normalisation (BN) is widely used in conventional deep neural network training to harmonise the input-output distributions for each batch of data. However, federated learning, a distributed learning paradigm, faces the challenge of dealing with non-independent and identically distributed data…

2025

One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image Fusion

CVPR 2025poster

Advanced image fusion methods mostly prioritise high-level missions, where task interaction struggles with semantic gaps, requiring complex bridging mechanisms. In contrast, we propose to leverage low-level vision tasks from digital photography fusion, allowing for effective feature interaction thro…

2025

One-Shot Reference-based Structure-Aware Image to Sketch Synthesis

AAAI 2025technical

Generating sketches that accurately reflect the content of reference images presents numerous challenges. Current methods either require paired training data or fail to accommodate a wider range and diversity of sketch styles. While pre-trained diffusion models have shown strong text-based control c…

Cited by 0SourcePDFScholar
2025

R-DTI: Drug Target Interaction Prediction Based on Second-Order Relevance Exploration

AAAI 2025technical

Drug Target Interaction (DTI) prediction has witnessed promising performance boosts accompanied by advanced multimodal feature extraction. However, existing approaches suffer from two main difficulties. First, the complex protein structures cannot be well represented by current protein-sequence-base…

2025

Self-Supervised Learning and Image-Prompt Fusion for AIGC Image Quality Assessment

ICASSP 2025accepted

With the rapid advancement of artificial intelligence, the field of Artificial Intelligence Generated Content (AIGC) has seen significant growth. As AI-generated images (AIGIs) become increasingly prevalent, the AIGC image quality assessment(AIGCIQA) has gained critical importance. However, traditio…

Cited by 0SourceScholar
2025

Stroke2Sketch: Harnessing Stroke Attributes for Training-Free Sketch Generation

ICCV 2025poster

Generating sketches guided by reference styles requires precise transfer of stroke attributes, such as line thickness, deformation, and texture sparsity, while preserving semantic structure and content fidelity. To this end, we propose Stroke2Sketch, a novel training-free framework that introduces c…

2025

Towards a General Attention Framework on Gyrovector Spaces for Matrix Manifolds

NeurIPS 2025poster

Deep neural networks operating on non-Euclidean geometries have recently demonstrated impressive performance across various machine-learning applications. Several studies have extended the attention mechanism to different manifolds. However, most existing non-Euclidean attention models are tailored…

Cited by 0SourceScholar
2025

Understanding Matrix Function Normalizations in Covariance Pooling through the Lens of Riemannian Geometry

ICLR 2025poster

Global Covariance Pooling (GCP) has been demonstrated to improve the performance of Deep Neural Networks (DNNs) by exploiting second-order statistics of high-level representations. GCP typically performs classification of the covariance matrices by applying matrix function normalization, such as mat…

2024

Can ChatGPT Serve as a Multi-Criteria Decision Maker? A Novel Approach to Supplier Evaluation

ICASSP 2024accepted

Multi-Criteria Decision Making (MCDM) has found extensive applications across various domains such as business, engineering, education, and academia, with supplier evaluation being a quintessential task among them. Traditional MCDM models typically gather quantitative and qualitative data through me…

Cited by 0SourceScholar
2024

Generative-Based Fusion Mechanism for Multi-Modal Tracking

AAAI 2024technical

Generative models (GMs) have received increasing research interest for their remarkable capacity to achieve comprehensive understanding. However, their potential application in the domain of multi-modal tracking has remained unexplored. In this context, we seek to uncover the potential of harnessing…

2024

NERF-GAZE: A Head-Eye Redirection Parametric Model for Gaze Estimation

ICASSP 2024accepted

Gaze estimation is a fundamental aspect of many visual tasks. However, the high cost of acquiring gaze datasets with 3D annotations hinders the optimization and application of gaze estimation models. In this work, we propose a novel Head-Eye redirection parametric model based on Neural Radiance Fiel…

Cited by 0SourceScholar
2024

RMLR: Extending Multinomial Logistic Regression into General Geometries

NeurIPS 2024poster

Riemannian neural networks, which extend deep learning techniques to Riemannian spaces, have gained significant attention in machine learning. To better classify the manifold-valued features, researchers have started extending Euclidean multinomial logistic regression (MLR) into Riemannian manifolds…

2022

Noise Learning for Text Classification: A Benchmark

COLING 2022main

Noise Learning is important in the task of text classification which depends on massive labeled data that could be error-prone. However, we find that noise learning in text classification is relatively underdeveloped: 1. many methods that have been proven effective in the image domain are not explor…

Cited by 11SourcePDFScholar
2020

Peeking into occluded joints: A novel framework for crowd pose estimation

ECCV 2020poster

Although occlusion widely exists in nature and remains a fundamental challenge for pose estimation, existing heatmap-based approaches suffer serious degradation on occlusions. Their intrinsic problem is that they directly localize the joints based on visual information; however, the invisible joints…