← Search

Yuxiang Hu

9 accepted papers

2026

Rethinking Flow and Diffusion Bridge Models for Speech Enhancement

AAAI 2026technical

Flow matching and diffusion bridge models have emerged as leading paradigms in generative speech enhancement, modeling stochastic processes between paired noisy and clean speech signals based on principles such as flow matching, score matching, and Schrödinger bridge. In this paper, we present a fra

Cited by 0SourcePDFScholar
2024

A Lightweight Hybrid Multi-Channel Speech Extraction System with Directional Voice Activity Detection

ICASSP 2024accepted

Although deep learning (DL) based end-to-end models have shown outstanding performance in multi-channel speech extraction, their practical applications on edge devices are restricted due to their high computational complexity. In this paper, we propose a hybrid system that can more effectively integ…

Cited by 0SourceScholar
2024

GTCRN: A Speech Enhancement Model Requiring Ultralow Computational Resources

ICASSP 2024accepted

While modern deep learning-based models have significantly outperformed traditional methods in the area of speech enhancement, they often necessitate a lot of parameters and extensive computational power, making them impractical to be deployed on edge devices in real-world applications. In this pape…

Cited by 0SourceScholar
2023

A Low-Latency Hybrid Multi-Channel Speech Enhancement System For Hearing Aids

ICASSP 2023accepted

This paper summarizes a hybrid multi-channel speech enhancement system for the ICASSP Signal Processing Grand Challenge: Clarity Challenge (Speech Enhancement for Hearing Aids) 2023. The system consists of a rule-based dereverberation module, a multi-channel enhancement module, and a post-processing…

Cited by 0SourceScholar
2023

Convolutional Recurrent MetriCGAN With Spectral Dimension Compression For Full-Band Speech Enhancement

ICASSP 2023accepted

MetricGAN and its variations have been proven to be an effective wide-band speech enhancement model. In this paper, we expand it to full-band enhancement by combining our recently proposed learnable spectral dimension compression mapping strategy. The encoder-decoder structure with a time-frequency…

Cited by 0SourceScholar
2021

BiToD: A Bilingual Multi-Domain Dataset For Task-Oriented Dialogue Modeling

NeurIPS 2021poster

Task-oriented dialogue (ToD) benchmarks provide an important avenue to measure progress and develop better conversational agents. However, existing datasets for end-to-end ToD modeling are limited to a single language, hindering the development of robust end-to-end ToD systems for multilingual count…

Cited by 58SourcecodeScholar
2020

DeepWeave: Accelerating Job Completion Time with Deep Reinforcement Learning-based Coflow Scheduling

IJCAI 2020poster

To improve the processing efficiency of jobs in distributed computing, the concept of coflow is proposed. A coflow is a collection of flows that are semantically correlated in a multi-stage computation task. A job consists of multiple coflows and can be usually formulated as a Directed-Acyclic Graph…

Cited by 0SourcePDFScholar
2020

Improving the Scalability of Deep Reinforcement Learning-Based Routing with Control on Partial Nodes

ICASSP 2020accepted

Machine Learning (ML)-based routing optimization has been proposed to optimize the performance of flow routing for future networks, such as Software-Defined Networks (SDNs). However, existing studies are either hard to converge for large networks or vulnerable to topology changes. In this paper, we…

Cited by 0SourceScholar
2020

QOS-Aware Flow Control for Power-Efficient Data Center Networks with Deep Reinforcement Learning

ICASSP 2020accepted

Reducing the power consumption and maintaining the Flow Completion Time (FCT) for the Quality of Service (QoS) of applications in Data Center Networks (DCNs) are two major concerns for data center operators. However, existing works either fail in guaranteeing the QoS due to the neglect of the FCT co…

Cited by 0SourceScholar