← Search

Xiao-Jun Wu

16 accepted papers

2026

FusionRegister: Every Infrared and Visible Image Fusion Deserves Registration

CVPR 2026

Spatial registration across different visual modalities is a critical but formidable step in multi-modality image fusion for real-world perception. Although several methods are proposed to address this issue, the existing registration-based fusion methods typically require extensive pre-registration

Cited by 0SourcecodeScholar
2026

Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language Models

CVPR 2026

The rapid progress of Multi-Modal Large Language Models (MLLMs) has significantly advanced downstream applications. However, this progress also exposes serious transferable adversarial vulnerabilities. In general, existing adversarial attacks against MLLMs typically rely on surrogate models trained

Cited by 0SourcecodeScholar
2026

Riemannian Graph Convolutional Network for Skeleton-Based Two-Person Interaction Recognition

IJCAI 2026

In the field of skeleton-based human action recognition, Graph Convolutional Networks (GCNs) have become a dominant framework. However, existing GCN-based approaches often treat the sequences of two-person interaction as separate entities, ignoring the inherent semantic dependencies and spatial corr

Cited by 0Scholar
2026

Wasserstein-Aligned Hyperbolic Multi-View Clustering

AAAI 2026technical

Multi-view clustering (MVC) aims to uncover the latent structure of multi-view data by learning view-common and view-specific information. Although recent studies have explored hyperbolic representations for better tackling the representation gap between different views, they focus primarily on inst

Cited by 0SourcePDFScholar
2025

A Correlation Manifold Self-Attention Network for EEG Decoding

IJCAI 2025

Riemannian neural networks, which generalize the deep learning paradigm to non-Euclidean geometries, have garnered widespread attention across diverse applications in artificial intelligence. Among these, the representative attention models have been studied on various non-Euclidean spaces to geomet

2025

Learning to Normalize on the SPD Manifold under Bures-Wasserstein Geometry

CVPR 2025poster

Covariance matrices have proven highly effective across many scientific fields. Since these matrices lie within the Symmetric Positive Definite (SPD) manifold--a Riemannian space with intrinsic non-Euclidean geometry, the primary challenge in representation learning is to respect this underlying geo…

2024

A Grassmannian Manifold Self-Attention Network for Signal Classification

IJCAI 2024poster

In the community of artificial intelligence, significant progress has been made in encoding sequential data using deep learning techniques. Nevertheless, how to effectively mine useful information from channel dimensions remains a major challenge, as these features have a submanifold structure. Line…

2024

Riemannian Multinomial Logistics Regression for SPD Neural Networks

CVPR 2024poster

Deep neural networks for learning Symmetric Positive Definite (SPD) matrices are gaining increasing attention in machine learning. Despite the significant progress most existing SPD networks use traditional Euclidean classifiers on an approximated space rather than intrinsic classifiers that accurat…

2024

SCD-Net: Spatiotemporal Clues Disentanglement Network for Self-Supervised Skeleton-Based Action Recognition

AAAI 2024technical

Contrastive learning has achieved great success in skeleton-based action recognition. However, most existing approaches encode the skeleton sequences as entangled spatiotemporal representations and confine the contrasts to the same level of representation. Instead, this paper introduces a novel cont…

2023

RGBD1K: A Large-Scale Dataset and Benchmark for RGB-D Object Tracking

AAAI 2023technical

RGB-D object tracking has attracted considerable attention recently, achieving promising performance thanks to the symbiosis between visual and depth channels. However, given a limited amount of annotated RGB-D tracking data, most state-of-the-art RGB-D trackers are simple extensions of high-perform…

2023

Riemannian Local Mechanism for SPD Neural Networks

AAAI 2023technical

The Symmetric Positive Definite (SPD) matrices have received wide attention for data representation in many scientific areas. Although there are many different attempts to develop effective deep architectures for data processing on the Riemannian manifold of SPD matrices, very few solutions explicit…

2022

Where to Attack: A Dynamic Locator Model for Backdoor Attack in Text Classifications

COLING 2022main

Nowadays, deep-learning based NLP models are usually trained with large-scale third-party data which can be easily injected with malicious backdoors. Thus, BackDoor Attack (BDA) study has become a trending research to help promote the robustness of an NLP system. Text-based BDA aims to train a poiso…

2019

Joint Group Feature Selection and Discriminative Filter Learning for Robust Visual Object Tracking

ICCV 2019poster

We propose a new Group Feature Selection method for Discriminative Correlation Filters (GFS-DCF) based visual object tracking. The key innovation of the proposed method is to perform group feature selection across both channel and spatial dimensions, thus to pinpoint the structural relevance of mult…

Cited by 242PDFcodeScholar
2018

Wing Loss for Robust Facial Landmark Localisation With Convolutional Neural Networks

CVPR 2018poster

We present a new loss function, namely Wing loss, for robust facial landmark localisation with Convolutional Neural Networks (CNNs). We first compare and analyse different loss functions including L2, L1 and smooth L1. The analysis of these loss functions suggests that, for the training of a CNN-bas…

Cited by 548SourcePDFScholar
2017

Dynamic Attention-Controlled Cascaded Shape Regression Exploiting Training Data Augmentation and Fuzzy-Set Sample Weighting

CVPR 2017poster

We present a new Cascaded Shape Regression (CSR) architecture, namely Dynamic Attention-Controlled CSR (DAC-CSR), for robust facial landmark detection on unconstrained faces. Our DAC-CSR divides facial landmark detection into three cascaded sub-tasks: face bounding box refinement, general CSR and at…

Cited by 120PDFScholar