← Search

Xian Wei

29 accepted papers

2026

Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates

AAAI 2026technical

Existing paper review methods often rely on superficial manuscript features or directly on large language models (LLMs), which are prone to hallucinations, biased scoring, and limited reasoning capabilities. Moreover, these methods often fail to capture the complex argumentative reasoning and negoti

Cited by 0SourcePDFScholar
2026

Curvature-Aware Captioning: Leveraging Geodesic Attention for 3D Scene Understanding

CVPR 2026

Accurate 3D scene description is fundamental to robotic navigation and augmented reality, yet current dense captioning methods face significant limitations in processing sparse point cloud data.Existing approaches that apply Euclidean embedding spaces struggle to simultaneously preserve fine-grained

Cited by 0SourceScholar
2026

The Sword, Shield, and Achilles’ Heel: Characterizing the Linguistic Inductive Bias of Large Language Models for Spatial Reasoning in Navigation Planning

IJCAI 2026

Large Language Model (LLM)-based navigation systems have commonly constructed expli cit spatial representations (e.g., topological graphs, semantic raster maps) and translated them into textual descriptions as LLMs’ inputs. However, the linguistic structures of such text-based spatial representation

Cited by 0Scholar
2025

3D Dense Captioning via Prototypical Momentum Distillation

ICRA 2025

3D dense captioning aims to describe the crucial regions in 3D visual scenes in the form of natural language. Recent prevailing approaches achieve promising results by leveraging complicated structures incorporated with large-scale models, which necessitate abundant parameters and pose challenges re

Cited by 0SourceScholar
2025

CE-FFT: Communication-Efficient Federated Fine-Tuning for Large Language Models via Quantization and In-Context Learning

ICASSP 2025accepted

Although Federated Fine-Tuning (FFT) facilitates the fine-tuning of Large Language Models (LLMs) across data owners without compromising their privacy, it suffers from severe communication overheads caused by numerous parameters of LLMs even with Parameter-Efficient Fine-Tuning (PEFT) methods. To ad…

Cited by 0SourceScholar
2025

Dual-BEV Nav: Dual-Layer BEV-Based Heuristic Path Planning for Robotic Navigation in Unstructured Outdoor Environments

ICRA 2025

Path planning with strong environmental adaptability plays a crucial role in robotic navigation in unstructured outdoor environments, especially in the case of low-quality location and map information. The path planning ability of a robot depends on the identification of the traversability of global

Cited by 2SourceScholar
2025

EqGAN: Reformation-based Feature Equalization Fusion for Few-shot Image Generation

ICASSP 2025accepted

Due to the absence or mismatch of semantic information, existing few-shot image generation methods suffer from unsatisfactory generation quality and diversity, which have minimal benefits as data augmentation for downstream classification tasks. Reformatting the contextual and textural information o…

Cited by 0SourceScholar
2025

FiTGAN: Content Fusion with Style Transformation for Few-shot Image Generation

ICASSP 2025accepted

Due to the semantic entanglement in fusion strategies or unstable training in complicated image transformations, existing few-shot image generation methods still suffer from low generation quality and diversity. To tackle the above problems, we propose a novel fusion- and transformation-based framew…

Cited by 0SourceScholar
2025

KiteRunner: Language-Driven Cooperative Local-Global Navigation Policy with UAV Mapping in Outdoor Environments

IROS 2025

Autonomous navigation in open-world outdoor environments faces challenges in integrating dynamic conditions, long-distance spatial reasoning, and semantic understanding. Traditional methods struggle to balance local planning, global planning, and semantic task execution, while existing large languag

Cited by 2SourceScholar
2025

PDDFormer: Pairwise Distance Distribution Graph Transformer for Crystal Material Property Prediction

IJCAI 2025

Crystal structures can be simplified as a periodic point set that repeats across three-dimensional space along an underlying lattice. Traditionally, crystal representation methods rely on descriptors such as lattice parameters, symmetry, and space groups to characterize the structure. However, in re

Cited by 0SourcePDFScholar
2025

R2Det: Exploring Relaxed Rotation Equivariance in 2D Object Detection

ICLR 2025poster

Group Equivariant Convolution (GConv) empowers models to explore underlying symmetry in data, improving performance. However, real-world scenarios often deviate from ideal symmetric systems caused by physical permutation, characterized by non-trivial actions of a symmetry group, resulting in asymmet…

2025

Relaxed Rotational Equivariance via G-Biases in Vision

AAAI 2025technical

Group Equivariant Convolution (GConv) can capture rotational equivariance from original data. It assumes uniform and strict rotational equivariance across all features as the transformations under the specific group. However, the presentation or distribution of real-world data rarely conforms to str…

2024

Exact Fusion via Feature Distribution Matching for Few-shot Image Generation

CVPR 2024poster

Few-shot image generation as an important yet challenging visual task still suffers from the trade-off between generation quality and diversity. According to the principle of feature-matching learning existing fusion-based methods usually fuse different features by using similarity measurements or a…

2024

Hyperbolic Graph Diffusion Model

AAAI 2024technical

Diffusion generative models (DMs) have achieved promising results in image and graph generation. However, real-world graphs, such as social networks, molecular graphs, and traffic graphs, generally share non-Euclidean topologies and hidden hierarchies. For example, the degree distributions of graphs…

2024

ProEqBEV: Product Group Equivariant BEV Network for 3D Object Detection in Road Scenes of Autonomous Driving

ICRA 2024poster

With the rapid development of autonomous driving systems, 3D object detection based on Bird’s Eye View (BEV) in road scenes has witnessed great progress over the past few years. As a road scene exhibits a part-whole hierarchy between the within objects and the scene itself, simple parts (e.g., roads…

Cited by 2SourceScholar
2024

SampDetox: Black-box Backdoor Defense via Perturbation-based Sample Detoxification

NeurIPS 2024poster

The advancement of Machine Learning has enabled the widespread deployment of Machine Learning as a Service (MLaaS) applications. However, the untrustworthy nature of third-party ML services poses backdoor threats. Existing defenses in MLaaS are limited by their reliance on training samples or white-…

Cited by 1SourcePDFScholar
2024

Social Lode: Human Trajectory Prediction with Latent Odes

ICASSP 2024accepted

Human trajectory prediction is crucial in human-computer interaction and even in the safety of autonomous driving. In this work, A new method, called Social Latent Ordinary Differential Equation (Social LODE), is introduced for predicting human trajectories. The backbone of Social LODE consists of a…

Cited by 0SourceScholar
2024

WaveAttack: Asymmetric Frequency Obfuscation-based Backdoor Attacks Against Deep Neural Networks

NeurIPS 2024poster

Due to the increasing popularity of Artificial Intelligence (AI), more and more backdoor attacks are designed to mislead Deep Neural Network (DNN) predictions by manipulating training samples or processes. Although backdoor attacks have been investigated in various scenarios, they still suffer from…

2023

DuEqNet: Dual-Equivariance Network in Outdoor 3D Object Detection for Autonomous Driving

ICRA 2023poster

Outdoor 3D object detection has played an essential role in the environment perception of autonomous driving. In complicated traffic situations, precise object recognition provides indispensable information for prediction and planning in the dynamic system, improving self-driving safety and reliabil…

Cited by 9SourceScholar
2022

Eliminating Backdoor Triggers for Deep Neural Networks Using Attention Relation Graph Distillation

IJCAI 2022poster

Due to the prosperity of Artificial Intelligence (AI) techniques, more and more backdoors are designed by adversaries to attack Deep Neural Networks (DNNs). Although the state-of-the-art method Neural Attention Distillation (NAD) can effectively erase backdoor triggers from DNNs, it still suffers fr…

2022

Geodesic Self-Attention for 3D Point Clouds

NeurIPS 2022accept

Due to the outstanding competence in capturing long-range relationships, self-attention mechanism has achieved remarkable progress in point cloud tasks. Nevertheless, point cloud object often has complex non-Euclidean spatial structures, with the behavior changing dynamically and unpredictably. Most…

Cited by 16SourcePDFScholar
2022

Learning Extremely Lightweight and Robust Model with Differentiable Constraints on Sparsity and Condition Number

ECCV 2022poster

"Learning lightweight and robust deep learning models is an enormous challenge for safety-critical devices with limited computing and memory resources, owing to robustness against adversarial attacks being proportional to network capacity. The community has extensively explored the integration of ad…

2022

Spatial-Frequency Domain Information Integration for Pan-Sharpening

ECCV 2022poster

"Pan-sharpening aims to generate the high-resolution multi-spectral (MS) images by fusing PAN images and low-resolution MS images. Despite the great advances, most existing pan-sharpening methods only work in the spatial domain and rarely explore the potential solution in frequency domain. In this p…

Cited by 102SourcePDFScholar
2016

Trace Quotient Meets Sparsity: A Method for Learning Low Dimensional Image Representations

CVPR 2016poster

This paper presents an algorithm that allows to learn low dimensional representations of images in an unsupervised manner. The core idea is to combine two criteria that play important roles in unsupervised representation learning, namely sparsity and trace quotient. The former is known to be a conve…

Cited by 20PDFScholar