← Search

Xiaohui Zhang

28 accepted papers

2026

Design and Implementation of EPM Based Modular Micro-UAVs for Autonomous Midair Docking

RA-L 2026

This letter presents a modular micro-UAV system based on electro-permanent magnet (EPM) technology, addressing the critical challenges of energy-constrained docking mechanisms in resource-limited micro aerial platforms. Built upon the Crazyflie 2.1 platform, our 75 g modular design (with battery) fe

Cited by 1SourceScholar
2026

End-to-End Diffusion-Based 3D Object Reconstruction From Robotic Tactile Sensing

RA-L 2026

Tactile sensing is essential for robotic perception in scenarios where visual input is limited or unavailable. In this work, we propose a fully tactile-based 3D object reconstruction framework that recovers object shapes exclusively from contact observations. A robotic system comprising a robotic ar

Cited by 1SourceScholar
2026

FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning

AAAI 2026technical

Representation learning is fundamental to modern machine learning, powering applications such as text retrieval and multimodal understanding. However, learning robust and generalizable representations remains challenging. While prior work has demonstrated that active noise injection, a form of data

Cited by 0SourcePDFScholar
2026

InstEmb: Instruction-Following Embeddings through Glimpses of the Future

ICML 2026poster

Recent advances have empowered large language models (LLMs) with remarkable fine-grained instruction-following capabilities in text generation tasks. However, embedding methods typically rely solely on the hidden state of the input's last token, limiting their ability to capture complete semantic si…

Cited by 0SourceScholar
2026

Voice-Driven Assistance and Resistance Modulation in a Soft Hip Exosuit Using a Transformer-Based Speech Recognition Model

ICRA 2026poster

Intuitive human–robot interfaces are essential to increase usability and personalization in wearable robotic assistive technologies. However, most current systems rely on pre-programmed or sensor-driven strategies that offer limited active user control online. To address this limitation, we present …

Cited by 0Scholar
2025

Adversarial Training and Gradient Optimization for Partially Deepfake Audio Localization

ICASSP 2025accepted

Partially deepfake audio localization is important in audio forensics. However, existing localization models for partially deepfake audio face two major challenges: distribution shifts between training and testing data as well as insufficient utilization of information from both manipulated regions…

Cited by 0SourceScholar
2025

Efficient Streaming LLM for Speech Recognition

ICASSP 2025accepted

Recent works have shown that prompting large language models with audio encodings can unlock speech recognition capabilities. However, existing techniques do not scale efficiently, especially while handling long form streaming audio inputs — not only do they extrapolate poorly beyond the audio lengt…

Cited by 0SourceScholar
2025

MMEditor: Multimodal Prompt-Driven 3D Gaussian Splatting Editing

ICASSP 2025accepted

We propose a multimodal 3D scene editing framework MMEditor to create or modify objects within an extant 3D Gaussian Splatting (3DGS) according to text and image prompts. MMEditor employs a multimodal image editing module to iteratively optimize 3D Gaussians in editing regions for delicate and multi…

Cited by 0SourceScholar
2025

Region-Based Optimization in Continual Learning for Audio Deepfake Detection

AAAI 2025technical

Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes when confronted with the diverse and evolving nature of real…

2025

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

ICLR 2025poster

The imitation of voice, targeted on specific speech attributes such as timbre and speaking style, is crucial in speech generation. However, existing methods rely heavily on annotated data, and struggle with effectively disentangling timbre and style, leading to challenges in achieving controllable g…

2024

Heterogeneous Multi-Robot Cooperation With Asynchronous Multi-Agent Reinforcement Learning

RA-L 2024

Multi-robot systems (MRSs) are becoming increasingly important in various domains. However, effective communication and coordination among multiple robots remain significant challenges. In this letter, we introduce a novel architecture for multi-robot decision-making and control based on multi-agent

Cited by 19SourceScholar
2024

Less Peaky and More Accurate CTC Forced Alignment by Label Priors

ICASSP 2024accepted

Connectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can cause inaccurate forced alignments (FA), especially at finer granularity, e.g., phoneme level. This paper aims at allevia…

Cited by 0SourceScholar
2024

Multi-Scale Permutation Entropy for Audio Deepfake Detection

ICASSP 2024accepted

With the widespread application of Automatic Speaker Verification (ASV) technology in security authentication, the threat of fake audio attacks looms as a malicious means compromising system security. In this study, we employ the multi-scale permutation entropy (MPE) in audio deepfake detection, whi…

Cited by 0SourceScholar
2024

Multimodal Representation Learning by Alternating Unimodal Adaptation

CVPR 2024poster

Multimodal learning which integrates data from diverse sensory modes plays a pivotal role in artificial intelligence. However existing multimodal learning methods often struggle with challenges where some modalities appear more dominant than others during multimodal learning resulting in suboptimal…

2024

What to Remember: Self-Adaptive Continual Learning for Audio Deepfake Detection

AAAI 2024technical

The rapid evolution of speech synthesis and voice conversion has raised substantial concerns due to the potential misuse of such technology, prompting a pressing need for effective audio deepfake detection mechanisms. Existing detection models have shown remarkable success in discriminating known de…

Cited by 29SourcePDFScholar
2023

Anchored Speech Recognition with Neural Transducers

ICASSP 2023accepted

Neural transducers have achieved human level performance on standard speech recognition benchmarks. However, their performance significantly degrades in the presence of cross-talk, especially when the primary speaker has a low signal-to-noise ratio. Anchored speech recognition refers to a class of m…

Cited by 2SourceScholar
2023

Do You Remember? Overcoming Catastrophic Forgetting for Fake Audio Detection

ICML 2023poster

Current fake audio detection algorithms have achieved promising performances on most datasets. However, their performance may be significantly degraded when dealing with audio of a different dataset. The orthogonal weight modification to overcome catastrophic forgetting does not consider the similar…

2023

Environment-Based Assistance Modulation for a Hip Exosuit via Computer Vision

RA-L 2023

Just like in humans vision plays a fundamental role in guiding adaptive locomotion, when designing the control strategy for a walking assistive technology, the use of computer vision may substantially improve modulation of the assistance based on the external environment. In this letter, we develope

Cited by 33SourceScholar
2023

Torchaudio-Squim: Reference-Less Speech Quality and Intelligibility Measures in Torchaudio

ICASSP 2023accepted

Measuring quality and intelligibility of a speech signal is usually a critical step in development of speech processing systems. To enable this, a variety of metrics to measure quality and intelligibility under different assumptions have been developed. Through this paper, we introduce tools and a s…

Cited by 124SourceScholar
2022

EMG-Driven Machine Learning Control of a Soft Glove for Grasping Assistance and Rehabilitation

RA-L 2022

In the field of rehabilitation robotics, transparent, precise and intuitive control of hand exoskeletons still represents a substantial challenge. In particular, the use of compliant systems often leads to a trade-off between lightness and material flexibility, and control precision. In this letter,

Cited by 49SourceScholar
2022

Enhancing Gait Assistance Control Robustness of a Hip Exosuit by Means of Machine Learning

RA-L 2022

Optimally synchronising the assistance provided by wearable devices with the human voluntary motion is still an open challenge in robotics. In order to provide accurate and robust assistance, this paper presents a novel approach that combines a layered implementation of a controller for an underactu

Cited by 16SourceScholar
2022

Omni-Sparsity DNN: Fast Sparsity Optimization for On-Device Streaming E2E ASR Via Supernet

ICASSP 2022accepted

From wearables to powerful smart devices, modern automatic speech recognition (ASR) models run on a variety of edge devices with different computational budgets. To navigate the Pareto front of model accuracy vs model size, researchers are trapped in a dilemma of optimizing model accuracy by trainin…

Cited by 0SourceScholar
2022

Streaming Transformer Transducer based Speech Recognition Using Non-Causal Convolution

ICASSP 2022accepted

This paper improves the streaming transformer transducer for speech recognition using non-causal convolution. Many works apply the causal convolution to improve streaming transformer ignoring the lookahead context. We propose to use non-causal convolution to process the center block and lookahead co…

Cited by 0SourceScholar
2022

Towards Measuring Fairness in Speech Recognition: Casual Conversations Dataset Transcriptions

ICASSP 2022accepted

The problem of machine learning systems demonstrating bias towards specific groups of individuals has been studied extensively, particularly in the Facial Recognition area, but much less so in Automatic Speech Recognition (ASR). This paper presents initial Speech Recognition results on “Casual Conve…

Cited by 53SourceScholar
2022

Underactuated Soft Hip Exosuit Based on Adaptive Oscillators to Assist Human Locomotion

RA-L 2022

Reproducing the mechanisms of human locomotion is a hard challenge. Assistive wearable devices in this context need to be lightweight, portable, and to adapt to the wearer’s walking pattern. Aiming to combine the aforementioned features, we developed a soft wearable exosuit to assist hip flexion dur

Cited by 49SourceScholar
2020

DEJA-VU: Double Feature Presentation and Iterated Loss in Deep Transformer Networks

ICASSP 2020accepted

Deep acoustic models typically receive features in the first layer of the network, and process increasingly abstract representations in the subsequent layers. Here, we propose to feed the input features at multiple depths in the acoustic model. As our motivation is to allow acoustic models to re-exa…

Cited by 0SourceScholar
2020

OOV Recovery with Efficient 2nd Pass Decoding and Open-vocabulary Word-level RNNLM Rescoring for Hybrid ASR

ICASSP 2020accepted

In this paper, we investigate out-of-vocabulary (OOV) word recovery in hybrid automatic speech recognition (ASR) systems, with emphasis on dynamic vocabulary expansion for both Weight Finite State Transducer (WFST)-based decoding and word-level RNNLM rescoring. We first describe our OOV candidate ge…

Cited by 0SourceScholar
2020

Transformer-Based Acoustic Modeling for Hybrid Speech Recognition

ICASSP 2020accepted

We propose and evaluate transformer-based acoustic models (AMs) for hybrid speech recognition. Several modeling choices are discussed in this work, including various positional embedding methods and an iterated loss to enable training deep transformers. We also present a preliminary study of using l…

Cited by 0SourceScholar