← Search

Xiangwei Zhu

5 accepted papers

2025

Multi-Grained Query-Guided Set Prediction Network for Grounded Multimodal Named Entity Recognition

AAAI 2025technical

Grounded Multimodal Named Entity Recognition (GMNER) is an emerging information extraction (IE) task, aiming to simultaneously extract entity spans, types, and corresponding visual regions of entities from given sentence-image pairs data. Recent unified methods employing machine reading comprehensio…

2025

Towards Efficient Image-goal Navigation: A Self-supervised Transformer-based Reinforcement Learning Approach

IROS 2025

Image-goal navigation is a crucial yet challenging task that requires an agent to navigate to a goal location specified by an image. Modular methods decompose the problem into distinct subtasks and often involve explicit map construction, which can struggle in complex, unstructured environments. In

Cited by 0SourcecodeScholar
2024

A New Representation of Universal Successor Features for Enhancing the Generalization of Target-Driven Visual Navigation

RA-L 2024

Target-driven visual navigation is a long-standing objective in the field of robotics. Deep reinforcement learning methods have demonstrated their effectiveness in developing target-driven visual navigation policies, yet they often struggle with generalization. Although extended reinforcement learni

Cited by 5SourceScholar
2024

Acoustic-VINS: Tightly Coupled Acoustic-Visual-Inertial Navigation System for Autonomous Underwater Vehicles

RA-L 2024

In this work, we present an acoustic-visual-inertial navigation system (Acoustic-VINS) for underwater robot localization. Specifically, we address the problem of the global position of the underwater visual-inertial navigation system being inappreciable by tightly coupling the long baseline (LBL) sy

Cited by 17SourceScholar
2024

Parsing All Adverse Scenes: Severity-Aware Semantic Segmentation with Mask-Enhanced Cross-Domain Consistency

AAAI 2024technical

Although recent methods in Unsupervised Domain Adaptation (UDA) have achieved success in segmenting rainy or snowy scenes by improving consistency, they face limitations when dealing with more challenging scenarios like foggy and night scenes. We argue that these prior methods excessively focus on w…

Cited by 8SourcePDFScholar