← Search

Wen Yang

40 accepted papers

2026

BFMPF-Net: Bidirectional Frequency-Domain Modulation Progressive Fusion Network for Road Crack Segmentation

ICRA 2026poster

Recently, deep learning–based methods for road crack segmentation have achieved promising performance, particularly in robotic vision applications such as automated inspection and maintenance. However, most frequency-domain methods employ a decoupled processing strategy, overlooking the dynamic modu…

Cited by 0Scholar
2026

Dr.Occ: Depth- and Region-Guided 3D Occupancy from Surround-View Cameras for Autonomous Driving

CVPR 2026

3D semantic occupancy prediction is crucial for autonomous driving perception, offering comprehensive geometric scene understanding and semantic recognition. However, existing methods struggle with geometric misalignment in view transformation due to lack of pixel-level accurate depth estimation, an

Cited by 0SourceScholar
2026

From Contrast to Consistency: Rethinking Event-based Continuous-Time Optical Flow Estimation

CVPR 2026

Estimating continuous optical flow is a fundamental yet challenging problem in dynamic visual perception. Event-based cameras, with microsecond latency and high dynamic range, capture brightness changes asynchronously, offering a unique opportunity to model motion with fine temporal precision. Howev

Cited by 0SourceScholar
2026

GUI-ARP: ENHANCING GROUNDING WITH ADAPTIVE REGION PERCEPTION FOR GUI AGENTS

ICASSP 2026poster

Existing GUI grounding methods often struggle with fine-grained localization in high-resolution screenshots. To address this, we propose GUI-ARP, a novel framework that enables adaptive multi-stage inference. Equipped with the proposed Adaptive Region Perception (ARP) and Adaptive Stage Controlling…

Cited by 0SourcePDFScholar
2026

I2D-LocX: An Efficient, Precise and Robust Method for Camera Localization in LiDAR Maps

ICRA 2026poster

Camera localization within LiDAR maps has gained significant attention due to its potential for accurate positioning with low-cost and lightweight sensors compared to LiDAR-based systems. However, existing methods often prioritize localization accuracy, sometimes compromising efficiency, which can l…

Cited by 0SourceScholar
2026

Low-Rank and Sparsity Are All You Need: Exploring Robust Hierarchical Latent Subspaces for Transferable Adversarial Attack

ICML 2026poster

Adversarial examples pose serious threats to deep neural networks (DNNs), revealing fundamental vulnerabilities in model robustness. However, most existing adversarial attacks directly manipulate densely activated and highly redundant feature representations, which often leads to overfitting on surr…

Cited by 0SourceScholar
2026

SGE-GLoc: Semantic Gaussian Ellipsoid Scene Graphs for Efficient LiDAR Global Localization

RA-L 2026

Global localization, encompassing robust place recognition and precise transformation estimation, is crucial for mobile robot navigation when the global navigation satellite system (GNSS) is unavailable. While LiDAR-based approaches are favored for their accuracy in 3D perception and resilience to i

Cited by 0SourceScholar
2026

TwinTrack: Bridging Vision and Contact Physics for Real-Time Tracking of Unknown Objects in Contact-Rich Scenes

ICRA 2026poster

Real-time tracking of previously unseen, highly dynamic objects in contact-rich scenes, such as during dexterous in-hand manipulation, remains a major challenge. Pure vision-based approaches often fail under heavy occlusions due to frequent contact interactions and motion blur caused by abrupt impac…

2025

Asymmetric Hierarchical Difference-aware Interaction Network for Event-guided Motion Deblurring

AAAI 2025technical

Event cameras are bio-inspired sensors that are capable of capturing motion information with high temporal resolution, which show potential in aiding image motion deblurring recently. Most existing methods indiscriminately handle feature fusion of two modalities with symmetric unidirectional/bidirec…

2025

Differential-Flatness-Based Tracking Control for Tractor-Trailers in Reversing Maneuvers

IROS 2025

In this paper, we propose a differential-flatness-based controller (DFBC) for precise trajectory tracking of tractor-trailers, particularly during reversing maneuvers, which are challenging due to unstable equilibrium points. The proposed controller leverages the differential flatness property of tr

Cited by 0SourceScholar
2025

I2D-LocX: An Efficient, Precise and Robust Method for Camera Localization in LiDAR Maps

RA-L 2025

Camera localization within LiDAR maps has gained significant attention due to its potential for accurate positioning with low-cost and lightweight sensors compared to LiDAR-based systems. However, existing methods often prioritize localization accuracy, sometimes compromising efficiency, which can l

Cited by 2SourceScholar
2025

Implicit Cross-Lingual Rewarding for Efficient Multilingual Preference Alignment

ACL 2025finding

Direct Preference Optimization (DPO) has become a prominent method for aligning Large Language Models (LLMs) with human preferences. While DPO has enabled significant progress in aligning English LLMs, multilingual preference alignment is hampered by data scarcity. To address this, we propose a nove…

2025

KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical Reasoning

NeurIPS 2025poster

Recent advances have demonstrated that integrating reinforcement learning with rule-based rewards can significantly enhance the reasoning capabilities of large language models (LLMs), even without supervised fine-tuning (SFT). However, prevalent reinforcement learning algorithms such as GRPO and its…

Cited by 0SourcecodeScholar
2025

Language Imbalance Driven Rewarding for Multilingual Self-improving

ICLR 2025poster

Large Language Models (LLMs) have achieved state-of-the-art performance across numerous tasks. However, these advancements have predominantly benefited "first-class" languages such as English and Chinese, leaving many other languages underrepresented. This imbalance, while limiting broader applicati…

2025

Robust Visual Odometry Using Rigidly-Bundled Arbitrarily-Arranged Multi-Cameras

RA-L 2025

Making multi-camera visual SLAM systems easier to set up and more robust to the environment is attractive for vision robots. Existing monocular and binocular vision SLAM systems have narrow sensing Field-of-View (FoV), resulting in degenerated accuracy and limited robustness in textureless environme

Cited by 0SourcecodeScholar
2025

SMP-Attack: Boosting the Transferability of Feature Importance-based Adversarial Attack with Semantics-aware Multi-granularity Patchout

ICCV 2025poster

Transfer-based attacks pose a significant security threat to deep neural networks (DNNs), due to their strong performance on unseen models in real-world black-box scenarios. Building on this, feature importance-based attacks further improve the transferability of adversarial examples by effectively…

2025

Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual Questions

ACL 2025long

In visual question answering (VQA) context, users often pose ambiguous questions to visual language models (VLMs) due to varying expression habits. Existing research addresses such ambiguities primarily by rephrasing questions. These approaches neglect the inherently interactive nature of user inter…

2024

ConGeo: Robust Cross-view Geo-localization across Ground View Variations

ECCV 2024poster

"∗ Equal contribution Corresponding author (xuchangeis@whu.edu.cn) Cross-view geo-localization aims at localizing a ground-level query image by matching it to its corresponding geo-referenced aerial view. In real-world scenarios, the task requires accommodating diverse ground images captured by user…

Cited by 7SourcePDFScholar
2024

FE-DeTr: Keypoint Detection and Tracking in Low-quality Image Frames with Events

ICRA 2024poster

Keypoint detection and tracking in traditional image frames are often compromised by image quality issues such as motion blur and extreme lighting conditions. Event cameras offer potential solutions to these challenges by virtue of their high temporal resolution and high dynamic range. However, they…

Cited by 4SourcecodeScholar
2024

I2D-Loc++: Camera Pose Tracking in LiDAR Maps With Multi-View Motion Flows

RA-L 2024

Camera localization in LiDAR maps has become increasingly popular due to its promising ability to handle complex scenarios, surpassing the limitations of visual-only localization methods. However, existing approaches mostly focus on addressing the cross-modal 2D–3D gaps while overlooking the relatio

Cited by 4SourceScholar
2024

Motion Deblurring via Spatial-Temporal Collaboration of Frames and Events

AAAI 2024technical

Motion deblurring can be advanced by exploiting informative features from supplementary sensors such as event cameras, which can capture rich motion information asynchronously with high temporal resolution. Existing event-based motion deblurring methods neither consider the modality redundancy in sp…

2024

QuadricsNet: Learning Concise Representation for Geometric Primitives in Point Clouds

ICRA 2024poster

This paper presents a novel framework to learn a concise geometric primitive representation for 3D point clouds. Different from representing each type of primitive individually, we focus on the challenging problem of how to achieve a concise and uniform representation robustly. We employ quadrics to…

Cited by 5SourcecodeScholar
2024

Toward Robust Keypoint Detection and Tracking: A Fusion Approach With Event-Aligned Image Features

RA-L 2024

Robust keypoint detection and tracking are crucial for various robotic tasks. However, conventional cameras struggle under rapid motion and lighting changes, hindering local and edge feature extraction essential for keypoint detection and tracking. Event cameras offer advantages in such scenarios du

Cited by 12SourceScholar
2024

X-Instruction: Aligning Language Model in Low-resource Languages with Self-curated Cross-lingual Instructions

ACL 2024findings

Large language models respond well in high-resource languages like English but struggle in low-resource languages. It may arise from the lack of high-quality instruction following data in these languages. Directly translating English samples into these languages can be a solution but unreliable, lea…

2023

Dynamic Coarse-To-Fine Learning for Oriented Tiny Object Detection

CVPR 2023poster

Detecting arbitrarily oriented tiny objects poses intense challenges to existing detectors, especially for label assignment. Despite the exploration of adaptive label assignment in recent oriented object detectors, the extreme geometry shape and limited feature of oriented tiny objects still induce…

2023

Generalizing Event-Based Motion Deblurring in Real-World Scenarios

ICCV 2023poster

Event-based motion deblurring has shown promising results by exploiting low-latency events. However, current approaches are limited in their practical usage, as they assume the same spatial resolution of inputs and specific blurriness distributions. This work addresses these limitations and aims to…

Cited by 29PDFcodeScholar
2022

Lidar With Velocity: Correcting Moving Objects Point Cloud Distortion From Oscillating Scanning Lidars by Fusion With Camera

RA-L 2022

Lidar point cloud distortion from moving object is an important problem in autonomous driving, and recently becomes more demanding with the emerging of oscillating type lidars, which feature back-and-forth scanning patterns and complex distortions. Accurately correcting the point cloud distortion wo

Cited by 30SourcecodeScholar
2022

RFLA: Gaussian Receptive Field Based Label Assignment for Tiny Object Detection

ECCV 2022poster

"Detecting tiny objects is one of the main obstacles hindering the development of object detection. The performance of generic object detectors tends to drastically deteriorate on tiny object detection tasks. In this paper, we point out that either box prior in the anchor-based detector or point pri…

2022

Towards Robust Visual-Inertial Odometry with Multiple Non-Overlapping Monocular Cameras

IROS 2022poster

We present a Visual-Inertial Odometry (VIO) algorithm with multiple non-overlapping monocular cameras aiming at improving the robustness of the VIO algorithm. An initialization scheme and tightly-coupled bundle adjustment for multiple non-overlapping monocular cameras are proposed. With more stable…

Cited by 7SourceScholar
2022

Understanding the Dynamics of DNNs Using Graph Modularity

ECCV 2022poster

"There are good arguments to support the claim that deep neural networks (DNNs) capture better feature representations than the previous hand-crafted feature engineering, which leads to a significant performance improvement. In this paper, we move a tiny step towards understanding the dynamics of fe…

2021

Event-Based Synthetic Aperture Imaging With a Hybrid Network

CVPR 2021poster

Synthetic aperture imaging (SAI) is able to achieve the see through effect by blurring out the off-focus foreground occlusions and reconstructing the in-focus occluded targets from multi-view images. However, very dense occlusions and extreme lighting conditions may bring significant disturbances to…

Cited by 39PDFcodeScholar
2020

Efficient Spatial-Temporal Normalization of SAE Representation for Event Camera

RA-L 2020

Event-based cameras are a new type of vision sensor that can encode spatial-temporal context in a pixel-level event stream. Its appealing properties offer great potential for applications requiring low processing latency and low power consumption. As an effective representation of events, the surfac

Cited by 13SourceScholar
2020

Monocular Camera Localization in Prior LiDAR Maps with 2D-3D Line Correspondences

IROS 2020poster

Light-weight camera localization in existing maps is essential for vision-based navigation. Currently, visual and visual-inertial odometry (VO&VIO) techniques are well-developed for state estimation but with inevitable accumulated drifts and pose jumps upon loop closure. To overcome these problems,…

Cited by 64SourcecodeScholar