← Search

Wenhao Guan

9 accepted papers

2025

Dynamic Language Group-based MoE: Enhancing Code-Switching Speech Recognition with Hierarchical Routing

ICASSP 2025accepted

The Mixture of Experts (MoE) model is a promising approach for handling code-switching speech recognition (CS-ASR) tasks. However, the existing CS-ASR work on MoE has yet to leverage the advantages of MoE’s parameter scaling ability fully. This work proposes DLG-MoE, a Dynamic Language Group-based M…

Cited by 0SourceScholar
2025

SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow

ICASSP 2025accepted

Recently, flow matching based speech synthesis has significantly enhanced the quality of synthesized speech while reducing the number of inference steps. In this paper, we introduce SlimSpeech, a lightweight and efficient speech synthesis system based on rectified flow. We have built upon the existi…

Cited by 4SourceScholar
2025

Topo-Field: Topometric Mapping With Brain-Inspired Hierarchical Layout-Object-Position Fields

RA-L 2025

Mobile robots require comprehensive scene understanding to operate effectively in diverse environments, enriched with contextual information such as layouts, objects, and their relationships. Although advances like neural radiance fields (NeRFs) offer high-fidelity 3D reconstructions, they are compu

Cited by 3SourcecodeScholar
2024

FastOcc: Accelerating 3D Occupancy Prediction by Fusing the 2D Bird’s-Eye View and Perspective View

ICRA 2024poster

In autonomous driving, 3D occupancy prediction outputs voxel-wise status and semantic labels for more comprehensive understandings of 3D scenes compared with traditional perception tasks, such as 3D object detection and bird’s-eye view (BEV) semantic segmentation. Recent researchers have extensively…

Cited by 35SourceScholar
2024

Improving Multi-Speaker ASR With Overlap-Aware Encoding And Monotonic Attention

ICASSP 2024accepted

End-to-end (E2E) multi-speaker speech recognition with the serialized output training (SOT) strategy demonstrates good performance in modeling diverse speaker scenarios. However, the E2E architecture doesn’t explicitly address the modeling of overlapping speech areas, potentially limiting the model’…

Cited by 0SourceScholar
2024

MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech Synthesis

AAAI 2024technical

The style transfer task in Text-to-Speech (TTS) refers to the process of transferring style information into text content to generate corresponding speech with a specific style. However, most existing style transfer approaches are either based on fixed emotional labels or reference speech clips, whi…

2024

Multivariate Fourier Distribution Perturbation: Domain Shifts with Uncertainty in Frequency Domain

ICASSP 2024accepted

Diversifying training data techniques have achieved tremendous success in Domain Generalization (DG) tasks. The key to diversifying domain data is by increasing the types of domain styles. After investigating this issue from the perspective of the Fourier transform, the domain cue is found to be imp…

Cited by 0SourceScholar
2024

Reflow-TTS: A Rectified Flow Model for High-Fidelity Text-to-Speech

ICASSP 2024accepted

The diffusion models including Denoising Diffusion Probabilistic Models (DDPM) and score-based generative models have demonstrated excellent performance in speech synthesis tasks. However, its effectiveness comes at the cost of numerous sampling steps, resulting in prolonged sampling time required t…

Cited by 0SourceScholar
2024

SR-HuBERT : An Efficient Pre-Trained Model for Speaker Verification

ICASSP 2024accepted

Recently, pre-trained models (PTMs) have been extensively applied in speaker verification (SV) and greatly boosted system performance. However, mainstream PTMs currently concentrate on using frame-level universal representations. In this paper, we propose a novel pre-training framework that jointly…

Cited by 0SourceScholar