← Search

Xing Wei

51 accepted papers

2026

Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing

CVPR 2026

Document parsing is a fine-grained task where image resolution significantly impacts performance. While advanced research leveraging vision-language models benefits from high-resolution input to boost model performance, this often leads to a quadratic increase in the number of vision tokens and sign

Cited by 2SourcecodeScholar
2026

Dense Monocular SLAM in Real-Time with Structured Gaussian Representation

ICRA 2026poster

Monocular dense SLAM faces significant challenges in low-texture environments and under rapid camera motions. The recent development of 3D Gaussian Splatting (3DGS) offers a promising approach for real-time dense 3D reconstruction. However, existing 3DGS-based SLAM systems employ end-to-end optimiza…

Cited by 0SourceScholar
2026

Edge Self-Adversarial Augmentation Enhances Graph Contrastive Learning Against Neighborhood Inconsistency

AAAI 2026technical

Recent studies have shown that unsupervised graph contrastive learning (GCL) is vulnerable to adversarial attacks. Automatic adversarial augmentation techniques are proposed to improve both the effectiveness and robustness of GCL. Existing methods typically regard unsupervised contrastive loss as th

Cited by 0SourcePDFScholar
2026

FARTrack: Fast Autoregressive Visual Tracking with High Performance

ICLR 2026poster

Inference speed and tracking performance are two critical evaluation metrics in the field of visual tracking. However, high-performance trackers often suffer from slow processing speeds, making them impractical for deployment on resource-constrained devices. To alleviate this issue, we propose $\tex…

Cited by 0SourcecodeScholar
2026

Fast Mixture of Curvature-Aware Experts for Diverse and Dynamic Graph Topologies

ICML 2026poster

Dynamic graph learning, which focuses on modeling the merging, vanishing, and reconnection of nodes and edges, is crucial for real-world applications. In dynamic graphs, node neighborhoods often exhibit diverse and time-evolving topologies, including hierarchical, grid-like, and cyclic patterns. Exi…

Cited by 0SourceScholar
2026

GeoMind: Explicit Spatial Reasoning via Dual-Reference Geometric Modeling

IJCAI 2026

While Vision-Language Models (VLMs) excel at semantic understanding, they struggle to comprehend 3D spatial relationships from limited views. Their reliance on implicit geometric encoding often leads to severe hallucinations and inconsistencies in spatial reasoning tasks. To address this, we introdu

Cited by 0Scholar
2026

JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation

ICLR 2026poster

Vision-and-Language Navigation (VLN) requires an embodied agent to navigate through unseen environments, guided by natural language instructions and a continuous video stream. Recent advances in VLN have been driven by the powerful semantic understanding of Multimodal Large Language Models (MLLMs).…

Cited by 0SourcecodeScholar
2026

Omni-AD: A Large-scale and Versatile Benchmark for Industrial Anomaly Detection

CVPR 2026

Industrial Anomaly Detection (IAD) has attracted significant attention and witnessed rapid development. However, the advancement in this field is hindered by two key issues: the performance saturation of existing benchmarks, limiting discriminative evaluation of different IAD methods, and the absenc

Cited by 0SourceScholar
2026

Persistent Autoregressive Mapping with Traffic Rules for Autonomous Driving

AAAI 2026technical

Safe autonomous driving requires both accurate HD map construction and persistent awareness of traffic rules, even when their associated signs are no longer visible. However, existing methods either focus solely on geometric elements or treat rules as temporary classifications, failing to capture th

Cited by 0SourcePDFScholar
2026

PriorDrive: Enhancing Online HD Mapping with Unified Vector Priors

AAAI 2026technical

High-Definition Maps (HD maps) are essential for the precise navigation and decision-making of autonomous vehicles, yet their creation and upkeep present significant cost and timeliness challenges. The online construction of HD maps using on-board sensors has emerged as a promising solution; however

Cited by 0SourcePDFScholar
2026

TACOcc: Target-Adaptive Cross-Modal Fusion with Sequential Volume Rendering for 3D Semantic Occupancy Prediction

ICRA 2026poster

Multi-modal 3D semantic occupancy prediction remains challenged by two fundamental issues: (i) geometric--semantic misalignment introduced by fixed-neighborhood fusion under heterogeneous sensing distributions, and (ii) feature degradation with prediction inconsistency in dynamic scenes caused by sp…

Cited by 0Scholar
2025

Dense Monocular SLAM in Real-Time With Structured Gaussian Representation

RA-L 2025

Monocular dense SLAM faces significant challenges in low-texture environments and under rapid camera motions. The recent development of 3D Gaussian Splatting (3DGS) offers a promising approach for real-time dense 3D reconstruction. However, existing 3DGS-based SLAM systems employ end-to-end optimiza

Cited by 3SourceScholar
2025

Driving by the Rules: A Benchmark for Integrating Traffic Sign Regulations into Vectorized HD Map

CVPR 2025highlight

Ensuring adherence to traffic sign regulations is essential for both human and autonomous vehicle navigation. While current online mapping solutions often prioritize the construction of the geometric and connectivity layers of HD maps, overlooking the construction of the traffic regulation layer wit…

2025

Dynamic Integration of Task-Specific Adapters for Class Incremental Learning

CVPR 2025poster

Non-exemplar Class Incremental Learning (NECIL) enables models to continuously acquire new classes without retraining from scratch and storing old task exemplars, addressing privacy and storage issues. However, the absence of data from earlier tasks exacerbates the challenge of catastrophic forgetti…

Cited by 2SourcePDFScholar
2025

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training

CVPR 2025poster

Language-image pre-training faces significant challenges due to limited data in specific formats and the constrained capacities of text encoders. While prevailing methods attempt to address these issues through data augmentation and architecture modifications, they continue to struggle with processi…

2025

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

NeurIPS 2025spotlight

Vision–Language–Action (VLA) models are increasingly used for end-to-end driving due to their world knowledge and reasoning ability. Most prior work, however, inserts textual chains-of-thought (CoT) as intermediate steps tailored to the current scene. Such symbolic compressions can blur spatio-tempo…

Cited by 0SourcecodeScholar
2025

Learnable Fractional Reaction-Diffusion Dynamics for Under-Display ToF Imaging and Beyond

ICCV 2025poster

Under-display ToF imaging aims to achieve accurate depth sensing through a ToF camera placed beneath a screen panel. However, transparent OLED (TOLED) layers introduce severe degradations--such as signal attenuation, multi-path interference (MPI), and temporal noise--that significantly compromise de…

2025

MACA: Multi-Anchor Classification Approach for Unsupervised Domain Adaptation

ICASSP 2025accepted

Unsupervised Domain Adaptation for image classification aims to adapt models trained on a labeled source domain to an unlabeled target domain, improving target domain classification performance. However, previous UDA classification researches tend to assume the two domain distributions after domain…

Cited by 0SourceScholar
2025

Preference-driven Knowledge Distillation for Few-shot Node Classification

NeurIPS 2025poster

Graph neural networks (GNNs) can efficiently process text-attributed graphs (TAGs) due to their message-passing mechanisms, but their training heavily relies on the human-annotated labels. Moreover, the complex and diverse local topologies of nodes of real-world TAGs make it challenging for a single…

Cited by 0SourcecodeScholar
2025

SeqGrowGraph: Learning Lane Topology as a Chain of Graph Expansions

ICCV 2025poster

Accurate lane topology is essential for autonomous driving, yet traditional methods struggle to model the complex, non-linear structures--such as loops and bidirectional lanes--prevalent in real-world road structure. We present SeqGrowGraph, a novel framework that learns lane topology as a chain of…

Cited by 0SourcePDFScholar
2024

DYSON: Dynamic Feature Space Self-Organization for Online Task-Free Class Incremental Learning

CVPR 2024poster

In this paper we focus on a challenging Online Task-Free Class Incremental Learning (OTFCIL) problem. Different from the existing methods that continuously learn the feature space from data streams we propose a novel compute-and-align paradigm for the OTFCIL. It first computes an optimal geometry i.…

2024

Evolving Parameterized Prompt Memory for Continual Learning

AAAI 2024technical

Recent studies have demonstrated the potency of leveraging prompts in Transformers for continual learning (CL). Nevertheless, employing a discrete key-prompt bottleneck can lead to selection mismatches and inappropriate prompt associations during testing. Furthermore, this approach hinders adaptive…

2024

Neural Spectral Decomposition for Dataset Distillation

ECCV 2024poster

"In this paper, we propose Neural Spectrum Decomposition, a generic decomposition framework for dataset distillation. Unlike previous methods, we consider the entire dataset as a high-dimensional observation that is low-rank across all dimensions. We aim to discover the low-rank representation of th…

2024

Person-in-WiFi 3D: End-to-End Multi-Person 3D Pose Estimation with Wi-Fi

CVPR 2024poster

Wi-Fi signals in contrast to cameras offer privacy protection and occlusion resilience for some practical scenarios such as smart homes elderly care and virtual reality. Recent years have seen remarkable progress in the estimation of single-person 2D pose single-person 3D pose and multi-person 2D po…

Cited by 15SourcePDFScholar
2024

ReliaAvatar: A Robust Real-Time Avatar Animator with Integrated Motion Prediction

IJCAI 2024poster

Efficiently estimating the full-body pose with minimal wearable devices presents a worthwhile research direction. Despite significant advancements in this field, most current research neglects to explore full-body avatar estimation under low-quality signal conditions, which is prevalent in practica…

2024

Self-Training Domain Adaptation Via Weight Transmission Between Generators

ICASSP 2024accepted

Unsupervised domain adaptation (UDA) aims to transfer knowledge from the labeled source domain to the fully-unlabeled target domain, thus improving the classification performance of the target domain. Recently, self-training has shown its effectiveness on UDA. However, the feature space for generati…

Cited by 0SourceScholar
2023

Contrastive Domain Adaptation Via Delimitation Discriminator

ICASSP 2023accepted

Unsupervised domain adaptation aims to transfer the knowledge learned from the labeled source domain to the unlabeled target domain, thereby improving the classification performance of the target domain. Recent methods use contrastive learning to optimize this task, however, these methods only focus…

Cited by 5SourceScholar
2023

DKT: Diverse Knowledge Transfer Transformer for Class Incremental Learning

CVPR 2023poster

Deep neural networks suffer from catastrophic forgetting in class incremental learning, where the classification accuracy of old classes drastically deteriorates when the networks learn the knowledge of new classes. Many works have been proposed to solve the class incremental learning problem. Howev…

Cited by 14SourcePDFScholar
2023

Fast and Accurate Binary Neural Networks Based on Depth-Width Reshaping

AAAI 2023technical

Network binarization (i.e., binary neural networks, BNNs) can efficiently compress deep neural networks and accelerate model inference but cause severe accuracy degradation. Existing BNNs are mainly implemented based on the commonly used full-precision network backbones, and then the accuracy is imp…

2023

Knowledge Restore and Transfer for Multi-Label Class-Incremental Learning

ICCV 2023poster

Current class-incremental learning research mainly focuses on single-label classification tasks while multi-label class-incremental learning (MLCIL) with more practical application scenarios is rarely studied. Although there have been many anti-forgetting methods to solve the problem of catastrophic…

Cited by 18PDFcodeScholar
2023

Learning Symmetry-Aware Geometry Correspondences for 6D Object Pose Estimation

ICCV 2023poster

Current 6D pose estimation methods focus on handling objects that are previously trained, which limits their applications in real dynamic world. To this end, we propose a geometry correspondence-based framework, termed GCPose, to estimate 6D pose of arbitrary unseen objects without any re-training.…

Cited by 20PDFcodeScholar
2023

Sparse Parameterization for Epitomic Dataset Distillation

NeurIPS 2023poster

The success of deep learning relies heavily on large and diverse datasets, but the storage, preprocessing, and training of such data present significant challenges. To address these challenges, dataset distillation techniques have been proposed to obtain smaller synthetic datasets that capture the e…

2022

SOIT: Segmenting Objects with Instance-Aware Transformers

AAAI 2022technical

This paper presents an end-to-end instance segmentation framework, termed SOIT, that Segments Objects with Instance-aware Transformers. Inspired by DETR, our method views instance segmentation as a direct set prediction problem and effectively removes the need for many hand-crafted components like R…

2021

Direct Measure Matching for Crowd Counting

IJCAI 2021poster

Traditional crowd counting approaches usually use Gaussian assumption to generate pseudo density ground truth, which suffers from problems like inaccurate estimation of the Gaussian kernel sizes. In this paper, we propose a new measure-based counting approach to regress the predicted density maps to…

Cited by 48SourcePDFScholar
2021

Efficient Deep Image Denoising via Class Specific Convolution

AAAI 2021technical

Deep neural networks have been widely used in image denoising during the past few years. Even though they achieve great success on this problem, they are computationally inefficient which makes them inappropriate to be implemented in mobile devices. In this paper, we propose an efficient deep neural…

2021

Error-Aware Density Isomorphism Reconstruction for Unsupervised Cross-Domain Crowd Counting

AAAI 2021technical

This paper focuses on the unsupervised domain adaptation problem for video-based crowd counting, in which we use labeled data as source domain and unlabelled video data as target domain. It is challenging as there is a huge gap between the source and the target domain and no annotations of samples a…

2021

Few-Shot Class-Incremental Learning via Relation Knowledge Distillation

AAAI 2021technical

In this paper, we focus on the challenging few-shot class incremental learning (FSCIL) problem, which requires to transfer knowledge from old tasks to new ones and solves catastrophic forgetting. We propose the exemplar relation distillation incremental learning framework to balance the tasks of old…

Cited by 202SourcePDFScholar
2021

Learning to Count via Unbalanced Optimal Transport

AAAI 2021technical

Counting dense crowds through computer vision technology has attracted widespread attention. Most crowd counting datasets use point annotations. In this paper, we formulate crowd counting as a measure regression problem to minimize the distance between two measures with different supports and unequa…

Cited by 96SourcePDFScholar
2018

Grassmann Pooling as Compact Homogeneous Bilinear Pooling for Fine-Grained Visual Classification

ECCV 2018poster

Designing discriminative and invariant features is the key to visual recognition. Recently, the bilinear pooled feature matrix of Convolutional Neural Network (CNN) has shown to achieve state-of-the-art performance on a range of fine-grained visual recognition tasks. The bilinear feature matrix coll…

Cited by 121SourcePDFScholar