← Search

Xihong Wu

26 accepted papers

2026

Proactive Risk-Aware Trajectory Planning for Autonomous Driving in Unstructured Environments Via Reinforcement Learning with Adaptive Reward Design

ICRA 2026poster

Trajectory planning for autonomous driving in dynamic unstructured traffic remains a fundamental challenge. Existing methods are often reactive, i.e., they only respond to observed situations without explicitly anticipating future risks. Moreover, most reinforcement learning based approaches rely on…

Cited by 0Scholar
2025

A Novel Multimodal Method for Decoding Speech Perception from Brain Activities

ICASSP 2025accepted

Decoding speech from neural recordings has critical importance in application and scientific research. However, this task is still challenging with non-invasive recordings. Previous research has shown significant improvement in speech perception decoding task by leveraging wav2vec vectors and gives…

Cited by 0SourceScholar
2025

Cross-attention Inspired Selective State Space Models for Target Sound Extraction

ICASSP 2025accepted

The Transformer model, particularly its cross-attention module, is widely used for feature fusion in target sound extraction which extracts the signal of interest based on given clues. Despite its effectiveness, this approach suffers from low computational efficiency. Recent advancements in state sp…

Cited by 0SourceScholar
2025

DPGP: A Hybrid 2D-3D Dual Path Potential Ghost Probe Zone Prediction Framework for Safe Autonomous Driving

IROS 2025

Modern robots must coexist with humans in dense urban environments. A key challenge is the ghost probe problem, where pedestrians or objects unexpectedly rush into traffic paths. This issue affects both autonomous vehicles and human drivers. Existing works propose vehicle-to-everything (V2X) strateg

Cited by 2SourceScholar
2025

Online Iterative Learning with Forward Simulation for Sub-minimum End-effector Displacement Positioning

IROS 2025

Precision is a crucial performance indicator for robot arms. During interacting with human, high precision enables a robot arm to be used effectively and safely, while low precision may lead to safety issues. Traditional methods for improving robot arm precision rely on error compensation. However,

Cited by 0SourceScholar
2025

SILM: A Subjective Intent Based Low-Latency Framework for Multiple Traffic Participants Joint Trajectory Prediction

IROS 2025

Trajectory prediction is a fundamental technology for advanced autonomous driving systems and represents one of the most challenging problems in the field of cognitive intelligence. Accurately predicting the future trajectories of each traffic participant is a prerequisite for building high safety a

Cited by 0SourceScholar
2025

Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment

ICASSP 2025accepted

Auditory Attention Decoding (AAD) can help to determine the identity of the attended speaker during an auditory selective attention task, by analyzing and processing measurements of electroencephalography (EEG) data. Most studies on AAD are based on scalp-EEG signals in two-speaker scenarios, which…

Cited by 0SourceScholar
2024

A DenseNet-Based Method for Decoding Auditory Spatial Attention with EEG

ICASSP 2024accepted

Auditory spatial attention detection (ASAD) aims to decode the attended spatial location with EEG in a multiple-speaker setting. ASAD methods are inspired by the brain lateralization of cortical neural responses during the processing of auditory spatial attention, and show promising performance for…

Cited by 0SourceScholar
2024

A Hybrid Deep-Online Learning Based Method for Active Noise Control in Wave Domain

ICASSP 2024accepted

The traditional feedback Active Noise Control (ANC) algorithms are built upon linear filters, which leads to reduced performance when dealing with real-world noise. Deep learning-based feedback ANC algorithms have been proposed to overcome this problem. However, methods relying on pre-trained neural…

Cited by 0SourceScholar
2024

Semantic Reconstruction of Continuous Language from Meg Signals

ICASSP 2024accepted

Decoding language from neural signals holds considerable theoretical and practical importance. Previous research has indicated the feasibility of decoding text or speech from invasive neural signals. However, when using non-invasive neural signals, significant challenges are encountered due to their…

Cited by 0SourceScholar
2023

A Model-Based Hearing Compensation Method Using a Self-Supervised Framework

ICASSP 2023accepted

Hearing aids can improve auditory perception for hearing-impaired (HI) listeners, but even state-of-art devices provide only limited benefits if not configured correctly for the listeners. The prescriptive fittings of hearing aids ignore the individual difference among HI listeners with identical he…

Cited by 0SourceScholar
2023

TT-Net: Dual-Path Transformer Based Sound Field Translation in the Spherical Harmonic Domain

ICASSP 2023accepted

In the current method for the sound field translation tasks based on spherical harmonic (SH) analysis, the solution based on the additive theorem usually faces the problem of singular values caused by large matrix condition numbers. The influence of different distances and frequencies of the spheric…

Cited by 0SourceScholar
2020

Individual Distance-Dependent HRTFS Modeling Through A Few Anthropometric Measurements

ICASSP 2020accepted

The lack of data is a major problem in individual HRTF modeling. There are many HRTF databases, but each database only has limited HRTFs with different characteristics, such as distance-dependent HRTFs or individual HRTFs. How to effectively model HRTFs through several different databases is an impo…

Cited by 0SourceScholar
2020

Single-Channel Speech Separation Integrating Pitch Information Based on a Multi Task Learning Framework

ICASSP 2020accepted

Pitch is a critical cue for speech separation in humans' auditory perception. Although the technology of tracking pitch in single-talker speech succeeds in many applications, it's still a challenging problem to extract pitch information from speech mixtures in machine perception. In this paper, we a…

Cited by 0SourceScholar
2019

Improvements to the Matching Projection Decoding Method for Ambisonic System with Irregular Loudspeaker Layouts

ICASSP 2019accepted

The Ambisonic technique has been widely used for sound field recording and reproduction recently. However, the basic Ambisonic decoding method will break down when the playback loudspeakers distribute unevenly. Various methods have been proposed to solve this problem. This paper introduces several i…

Cited by 0SourceScholar
2019

Integrating Spectrotemporal Context into Features Based on Auditory Perception for Classification-based Speech Separation

ICASSP 2019accepted

Speech separation, which has been a challenging task for decades, especially at low signal-to-noise ratios (SNRs), can be cast as a classification problem. In such adverse acoustic environment, extracting robust features from noisy mixtures is crucial for successful classification. In the past studi…

Cited by 0SourceScholar
2018

A Time-Weighted Method for Predicting the Intelligibility of Speech in the Presence of Interfering Sounds

ICASSP 2018accepted

The speech intelligibility index (SII) has been widely used as an objective method of predicting speech intelligibility, but its traditional form is most effective predicting speech intelligibility scores under stationary noise but not more challenging conditions (e.g., competing noise interference)…

Cited by 0SourceScholar
2016

Biped robot falling motion control with human-inspired active compliance

IROS 2016poster

Protecting robot from broken of falling is always a challenge issue for a bipedal humanoid robot in dealing with various locomotion related tasks to serve human society, especially as the assigned tasks turns increasingly complicated and the corresponding real environment gets more and more complex.…

Cited by 13SourceScholar
2015

Constructing long short-term memory based deep recurrent neural networks for large vocabulary speech recognition

ICASSP 2015accepted

Long short-term memory (LSTM) based acoustic modeling methods have recently been shown to give state-of-the-art performance on some speech recognition tasks. To achieve a further performance improvement, in this research, deep extensions on LSTM are investigated considering that deep hierarchical mo…

Cited by 0SourceScholar
2015

Improving long short-term memory networks using maxout units for large vocabulary speech recognition

ICASSP 2015accepted

Long short-tem memory (LSTM) recurrent neural networks have been shown to give state-of-the-art performance on many speech recognition tasks. To achieve a further performance improvement, in this paper, maxout units are proposed to be integrated with the LSTM cells, considering those units have brou…

Cited by 0SourceScholar