← Search

Wei Xie

18 accepted papers

2025

NaviFormer: A Spatio-Temporal Context-Aware Transformer for Object Navigation

AAAI 2025technical

Learning discriminative state representations of agents, encompassing the spatial layout and temporal pose trajectory, is essential for effective navigation decisions. However, existing approaches often rely on simplistic plain networks for navigation information fusion, overlooking the complex long…

2025

PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric Network

NeurIPS 2025poster

With the exponential growth of video content, aiming at localizing relevant video moments based on natural language queries, video moment retrieval (VMR) has gained significant attention. Existing weakly supervised VMR methods focus on designing various feature modeling and modal interaction modules…

Cited by 0SourcecodeScholar
2025

Revisiting Generative Infrared and Visible Image Fusion Based on Human Cognitive Laws

NeurIPS 2025spotlight

Existing infrared and visible image fusion methods often face the dilemma of balancing modal information. Generative fusion methods reconstruct fused images by learning from data distributions, but their generative capabilities remain limited. Moreover, the lack of interpretability in modal informat…

Cited by 0SourceScholar
2024

Cross-Modal Multiscale Difference-Aware Network for Joint Moment Retrieval and Highlight Detection

ICASSP 2024accepted

Since the goals of both Moment Retrieval (MR) and Highlight Detection (HD) are to quickly obtain the required content from the video according to user needs, several works have attempted to take advantage of the commonality between both tasks to design transformer-based networks for joint MR and HD.…

Cited by 0SourceScholar
2024

TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection

AAAI 2024technical

Video moment retrieval (MR) and highlight detection (HD) based on natural language queries are two highly related tasks, which aim to obtain relevant moments within videos and highlight scores of each video clip. Recently, several methods have been devoted to building DETR-based networks to solve bo…

2023

Clean Sample Guided Self-Knowledge Distillation for Image Classification

ICASSP 2023accepted

For two-stage knowledge distillation, the combination with Data Augmentation (DA) is straightforward and effective. Yet, for online Self-knowledge Distillation (SD), DA is not always beneficial because of the absence of a trustworthy teacher model. To address this issue, this paper proposes an SD me…

Cited by 0SourceScholar
2023

Nonrigid Object Contact Estimation With Regional Unwrapping Transformer

ICCV 2023poster

Acquiring contact patterns between hands and nonrigid objects is a common concern in the vision and robotics community. However, existing learning-based methods focus more on contact with rigid ones from monocular images. When adopting them for nonrigid contact, a major problem is that the existing…

Cited by 3PDFScholar
2023

Reconstructing Interacting Hands with Interaction Prior from Monocular Images

ICCV 2023poster

Reconstructing interacting hands from monocular images is indispensable in AR/VR applications. Most existing solutions rely on the accurate localization of each skeleton joint. However, these methods tend to be unreliable due to the severe occlusion and confusing similarity among adjacent hand parts…

Cited by 26PDFcodeScholar
2023

Robust Target Interception Strategy for a USV With Experimental Validation

RA-L 2023

This letter addresses the problem of designing an interception strategy for an underactuated uncrewed surface vessel (USV) in the presence of uncertain external disturbances and unknown internal parameters i.e., linear and nonlinear damping coefficients, vehicle mass, etc. The interception strategy

Cited by 11SourceScholar
2022

Flexible Collision-free Platooning Method for Unmanned Surface Vehicle with Experimental Validations

IROS 2022poster

This paper addresses the flexible formation problem for unmanned surface vehicles in the presence of obstacles. Building upon the leader-follower formation scheme, a hybrid line-of-sight based flexible platooning method is proposed for follower vehicle to keep tracking the leader ship. A fusion arti…

Cited by 3SourceScholar
2020

Clustering Driven Deep Autoencoder for Video Anomaly Detection

ECCV 2020poster

Because of the ambiguous definition of anomaly and the complexity of real data, anomaly detection in videos is one of the most challenging problems in intelligent video surveillance. Since the abnormal events are usually different from normal events in appearance and/or in motion behavior, we addres…

Cited by 290SourcePDFScholar
2018

Robust Widely Widely Beamforming via the Technique of Shrinkage for Steering Vector Estimation

ICASSP 2018accepted

In this paper, two novel robust widely linear beamforming algorithms based on the technique of shrinkage are proposed, i.e., the WL-RBLW and the WL-OAS. Firstly, in order to remove the signal-of-interest's (SOl's) component from the sample covariance matrix (SCM), the augmented interference-plus-noi…

Cited by 0SourceScholar