← Search

Xinyuan Qian

23 accepted papers

2026

ANALYTIC INCREMENTAL LEARNING FOR SOUND SOURCE LOCALIZATION WITH IMBALANCE RECTIFICATION

ICASSP 2026poster

Sound source localization (SSL) demonstrates remarkable results in controlled settings but struggles in real-world deployment due to dual imbalance challenges: intra-task imbalance arising from long-tailed direction-of-arrival (DoA) distributions, and inter-task imbalance induced by cross-task skews…

Cited by 0SourcePDFScholar
2026

AV-SSAN: Audio-Visual Selective DOA Estimation Through Explicit Multi-Band Semantic-Spatial Alignment

AAAI 2026technical

Audio-visual sound source localization (AV-SSL) estimates the position of sound sources by fusing auditory and visual cues. Current AV-SSL methodologies typically require spatially-paired audio-visual data and cannot selectively localize specific target sources. To address these limitations, we intr

Cited by 0SourcePDFScholar
2026

BEYOND LIPS: INTEGRATING GESTURE AND LIP CUES FOR ROBUST AUDIO-VISUAL SPEAKER EXTRACTION

ICASSP 2026poster

Most audio-visual speaker extraction methods rely on synchronized lip recording to isolate the speech of a target speaker from a multi-talker mixture. However, in natural human communication, co-speech gestures are also temporally aligned with speech, often emphasizing specific words or syllables. T…

Cited by 0SourcePDFScholar
2026

MartDE: A Privacy-Preserving and Cost-Efficient Evaluation Framework for Data Marketplaces

AAAI 2026technical

The development of machine learning models increasingly relies on high-quality data that resides in private domains. To enable secure and value-driven data exchange under strict privacy regulations, federated learning (FL) has emerged as a key primitive by enabling the trading of model utilities ins

Cited by 0SourcePDFScholar
2026

PERFORMSINGER: MULTIMODAL SINGING VOICE SYNTHESIS LEVERAGING SYNCHRONIZED LIP CUES FROM SINGING PERFORMANCE VIDEOS

ICASSP 2026poster

Existing singing voice synthesis (SVS) models largely rely on fine-grained, phoneme-level durations, which limits their practical application. These methods overlook the complementary role of visual information in duration prediction.To address these issues, we propose PerformSinger, a pioneering mu…

Cited by 0SourcePDFScholar
2026

PTSE-T: PRESENTATION TARGET SPEAKER EXTRACTION USING UNALIGNED TEXT CUES

ICASSP 2026poster

Target Speaker Extraction (TSE) aims to extract the clean speech of the target speaker in an audio mixture, eliminating irrelevant background noise and speech. While prior work has explored various auxiliary cues including pre-recorded speech, visual information, and spatial information, the acquisi…

Cited by 0SourcePDFScholar
2025

Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition

ICASSP 2025accepted

Recently, end-to-end automatic speech recognition has become the mainstream approach in both industry and academia. To optimize system performance in specific scenarios, the Weighted Finite-State Transducer (WFST) is extensively used to integrate acoustic and language models, leveraging its capacity…

Cited by 0SourceScholar
2025

FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles

AAAI 2025technical

Humans can perceive speakers’ characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vision-driven Text-to-speech ( TTS ) scholars grounded their investigations on real-person faces, thereby restricting effec…

2025

M2PAIR: A High-Quality Acoustic Impulse Response Computation Model

ICASSP 2025accepted

Acoustic Impulse Response (AIR) provides crucial spatial information about the environment, significantly enhancing audio immersion. However, achieving high perceptual quality while computing AIR in real-time for interactive audio-video media (IAVM) presents a challenging problem. This study propose…

Cited by 0SourceScholar
2024

GLMB 3D Speaker Tracking with Video-Assisted Multi-Channel Audio Optimization Functions

ICASSP 2024accepted

Speaker tracking plays a significant role in numerous real-world human robot interaction (HRI) applications. In recent years, there has been a growing interest in utilizing multi-sensory information, such as complementary audio and visual signals, to address the challenges of speaker tracking. Despi…

Cited by 0SourceScholar
2024

LOCSELECT: Target Speaker Localization with an Auditory Selective Hearing Mechanism

ICASSP 2024accepted

The prevailing noise-resistant and reverberation-resistant localization algorithms primarily emphasize separating and providing directional output for each speaker in multi-speaker scenarios, without association with the identity of speakers. In this paper, we present a target speaker localization a…

Cited by 0SourceScholar
2023

A Miniaturised Camera-based Multi-Modal Tactile Sensor

ICRA 2023poster

In conjunction with huge recent progress in cam-era and computer vision technology, camera-based sensors have increasingly shown considerable promise in relation to tactile sensing. In comparison to competing technologies (be they resistive, capacitive or magnetic based), they offer super-high-resol…

Cited by 10SourceScholar
2023

L${3}$ F-TOUCH: A Wireless GelSight With Decoupled Tactile and Three-Axis Force Sensing

RA-L 2023

GelSight sensors that estimate contact geometry and force by reconstructing the deformation of their soft elastomer from images would yield poor force measurements when the elastomer deforms uniformly or reaches deformation saturation. Here we present an L <inline-formula xmlns:mml="http://www.w3.or

Cited by 37SourceScholar
2023

Ripple Sparse Self-Attention for Monaural Speech Enhancement

ICASSP 2023accepted

The use of Transformer represents a recent success in speech enhancement. However, as its core component, self-attention suffers from quadratic complexity, which is computationally prohibited for long speech recordings. Moreover, it allows each time frame to attend to all time frames, neglecting the…

Cited by 10SourceScholar
2023

Seeing What You Said: Talking Face Generation Guided by a Lip Reading Expert

CVPR 2023poster

Talking face generation, also known as speech-to-lip generation, reconstructs facial motions concerning lips given coherent speech input. The previous studies revealed the importance of lip-speech synchronization and visual quality. Despite much progress, they hardly focus on the content of lip move…

2023

Self-Convolution for Automatic Speech Recognition

ICASSP 2023accepted

Self-attention plays a significant role in recent automatic speech recognition (ASR) models with promising results. However, it suffers from high computational complexity and weak capability in modeling local information. In contrast, the convolutional neural network (CNN) is computationally effecti…

Cited by 0SourceScholar
2023

Stream Attention Based U-Net for L3DAS23 Challenge

ICASSP 2023accepted

Machine learning applications of 3D audio are gaining increasing interest in recent years. In this paper, we propose a stream attention based U-Net to remove background noise and reverberation based on ICASSP Signal Processing Grand Challenge 2023: L3DAS23 Challenge<sup xmlns:mml="http://www.w3.org/…

Cited by 0SourceScholar
2021

GCC-PHAT with Speech-oriented Attention for Robotic Sound Source Localization

ICRA 2021poster

Robotic audition is a basic sense that helps robots perceive the surroundings and interact with humans. Sound Source Localization (SSL) is an essential module for a robotic system. However, the performance of most sound source localization techniques degrades in noisy and reverberant environments du…

Cited by 19SourceScholar
2021

Multi-Target DoA Estimation with an Audio-Visual Fusion Mechanism

ICASSP 2021accepted

Most of the prior studies in the spatial Direction of Arrival (DoA) domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With this motivation, we propose to use neural networks with audio and visual signals for multi-speaker local…

Cited by 0SourceScholar
2019

Accurate Target Annotation in 3D from Multimodal Streams

ICASSP 2019accepted

Accurate annotation is fundamental to quantify the performance of multi-sensor and multi-modal object detectors and trackers. However, invasive or expensive instrumentation is needed to automatically generate these annotations. To mitigate this problem, we present a multi-modal approach that leverag…

Cited by 0SourceScholar
2018

3D Mouth Tracking from a Compact Microphone Array Co-Located with a camera

ICASSP 2018accepted

We address the 3D audio-visual mouth tracking problem when using a compact platform with co-located audio-visual sensors, without a depth camera. In particular, we propose a multi-modal particle filter that combines a face detector and 3D hypothesis mapping to the image plane. The audio likelihood c…

Cited by 0SourceScholar
2017

3D audio-visual speaker tracking with an adaptive particle filter

ICASSP 2017accepted

We propose an audio-visual fusion algorithm for 3D speaker tracking from a localised multi-modal sensor platform composed of a camera and a small microphone array. After extracting audio-visual cues from individual modalities we fuse them adaptively using their reliability in a particle filter frame…

Cited by 0SourceScholar