← Search

Cong Han

17 accepted papers

2025

Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation

ICASSP 2025accepted

Transformers have been the most successful architecture for various speech modeling tasks, including speech separation. However, the self-attention mechanism in transformers with quadratic complexity is inefficient in computation and memory. Recent models incorporate new layers and modules along wit…

Cited by 0SourceScholar
2025

Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis

ICASSP 2025accepted

It is too early to conclude that Mamba is a better alternative to transformers for speech before comparing Mamba with transformers in terms of both performance and efficiency in multiple speech-related tasks. To reach this conclusion, we propose and evaluate three models for three tasks: Mamba-TasNe…

Cited by 0SourceScholar
2025

StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion

NAACL 2025long

The rapid development of large-scale text-to-speech (TTS) models has led to significant advancements in modeling diverse speaker prosody and voices. However, these models often face issues such as slow inference speeds, reliance on complex pre-trained neural codec representations, and difficulties i…

2025

UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis

ICCV 2025poster

Text-to-image generation has transformed content creation, yet precise visual text rendering remains challenging for generative models due to blurred glyphs, semantic inconsistencies, and limited style controllability. Current methods typically employ pre-rendered glyph images as conditional inputs,…

Cited by 0SourcePDFScholar
2024

Exploring Self-supervised Contrastive Learning of Spatial Sound Event Representation

ICASSP 2024accepted

In this study, we present a simple multi-channel framework for contrastive learning (MC-SimCLR) to encode 'what' and 'where' of spatial audios. MC-SimCLR learns joint spectral and spatial representations from unlabeled spatial audios, thereby enhancing both event classification and sound localizatio…

Cited by 0SourceScholar
2023

Open-Vocabulary Semantic Segmentation with Decoupled One-Pass Network

ICCV 2023poster

Recently, the open-vocabulary semantic segmentation problem has attracted increasing attention and the best performing methods are based on two-stream networks: one stream for proposal mask generation and the other for segment classification using a pre-trained visual-language model. However, existi…

Cited by 48PDFcodeScholar
2023

Phoneme-Level Bert for Enhanced Prosody of Text-To-Speech with Grapheme Predictions

ICASSP 2023accepted

Large-scale pre-trained language models have been shown to be helpful in improving the naturalness of text-to-speech (TTS) models by enabling them to produce more naturalistic prosodic patterns. However, these models are usually word-level or sup-phoneme-level and jointly trained with phonemes, maki…

Cited by 0SourceScholar
2023

StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

NeurIPS 2023poster

In this paper, we present StyleTTS 2, a text-to-speech (TTS) model that leverages style diffusion and adversarial training with large speech language models (SLMs) to achieve human-level TTS synthesis. StyleTTS 2 differs from its predecessor by modeling styles as a latent random variable through dif…

2022

Improving Conversational Recommendation Systems’ Quality with Context-Aware Item Meta-Information

NAACL 2022findings

A key challenge of Conversational Recommendation Systems (CRS) is to integrate the recommendation function and the dialog generation function smoothly. Previous works employ graph neural networks with external knowledge graphs (KG) to model individual recommendation items and integrate KGs with lang…

2022

Multi-Channel Speech Denoising for Machine Ears

ICASSP 2022accepted

This work describes a speech denoising system for machine ears that aims to improve speech intelligibility and the overall listening experience in noisy environments. We recorded approximately 100 hours of audio data with reverberation and moderate environmental noise using a pair of microphone arra…

Cited by 0SourceScholar
2021

Dual-Path Modeling for Long Recording Speech Separation in Meetings

ICASSP 2021accepted

The continuous speech separation (CSS) is a task to separate the speech sources from a long, partially overlapped recording, which involves a varying number of speakers. A straightforward extension of conventional utterance-level speech separation to the CSS task is to segment the long recording wit…

Cited by 0SourceScholar
2021

Rethinking The Separation Layers In Speech Separation Networks

ICASSP 2021accepted

Modules in all existing speech separation networks can be categorized into single-input-multi-output (SIMO) modules and single-input-single-output (SISO) modules. SIMO modules generate more outputs than input, and SISO modules keep the numbers of input and output the same. While the majority of sepa…

Cited by 0SourceScholar