← Search

Lirong Dai

16 accepted papers

2025

Bridging Modality Gap with Large Speech and Language Models for End-to-End Speech-to-Text Translation

ICASSP 2025accepted

End-to-end speech-to-text translation (E2E ST) has increasingly aroused interest and attention recently, attempting to address the problem of data scarcity and modeling burden. Several attempts exploring the combination of Large Speech and Language Models into a unified model to improve E2E ST are c…

Cited by 0SourceScholar
2025

CSSinger: End-to-End Chunkwise Streaming Singing Voice Synthesis System Based on Conditional Variational Autoencoder

AAAI 2025technical

Singing Voice Synthesis (SVS) aims to generate singing voices of high fidelity and expressiveness. Conventional SVS systems usually utilize an acoustic model to transform a music score into acoustic features, followed by a vocoder to reconstruct the singing voice. It was recently shown that end-to-e…

2025

Leveraging Boolean Directivity Embedding for Binaural Target Speaker Extraction

ICASSP 2025accepted

Direction-based target speaker extraction (TSE) attracts a constant attention due to the convenience of direction acquisition over assistive video or enrollment audio. The direction clue heavily affects the TSE performance, which might be more seriously in the case of binaural setups due to the smal…

Cited by 0SourceScholar
2024

A Study of Multichannel Spatiotemporal Features and Knowledge Distillation on Robust Target Speaker Extraction

ICASSP 2024accepted

Target speaker extraction (TSE) based on direction of arrival (DOA) has a wide range of applications in e.g., remote conferencing, hearing aids, in-car speech interaction. Due to the inherent phase uncertainty, existing TSE methods usually suffer from speaker confusion within specific frequency band…

Cited by 0SourceScholar
2024

Adversarial Speech for Voice Privacy Protection from Personalized Speech Generation

ICASSP 2024accepted

The rapid progress in personalized speech generation technology, including personalized text-to-speech (TTS) and voice conversion (VC), poses a challenge in distinguishing between generated and real speech for human listeners, resulting in an urgent demand in protecting speakers' voices from malicio…

Cited by 12SourceScholar
2024

Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation

AAAI 2024technical

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is suffering from the scarcity of labeled multichannel data and complex ambient noise…

2024

Pre-Trained Acoustic-and-Textual Modeling for End-To-End Speech-To-Text Translation

ICASSP 2024accepted

End-to-end paradigm has aroused more and more interests and attention for improving speech-to-text translation (ST) recently. Existing end-to-end models mainly attributes and attempts to address the problem of modeling burden and data scarcity, while always fail to maintain both cross-modal and cros…

Cited by 0SourceScholar
2024

Sifisinger: A High-Fidelity End-to-End Singing Voice Synthesizer Based on Source-Filter Model

ICASSP 2024accepted

This paper presents an advanced end-to-end singing voice synthesis (SVS) system based on the source-filter mechanism that directly translates lyrical and melodic cues into expressive and high-fidelity human-like singing. Similarly to VISinger 2, the proposed system also utilizes training paradigms e…

Cited by 0SourceScholar
2023

Speech4Mesh: Speech-Assisted Monocular 3D Facial Reconstruction for Speech-Driven 3D Facial Animation

ICCV 2023poster

Recent audio2mesh-based methods have shown promising prospects for speech-driven 3D facial animation tasks. However, some intractable challenges are urgent to be settled. For example, the data-scarcity problem is intrinsically inevitable due to the difficulty of 4D data collection. Besides, current…

Cited by 10PDFScholar
2022

SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-training

EMNLP 2022main

The rapid development of single-modal pre-training has prompted researchers to pay more attention to cross-modal pre-training methods. In this paper, we propose a unified-modal speech-unit-text pre-training model, SpeechUT, to connect the representations of a speech encoder and a text decoder with a…

2021

TaLNet: Voice Reconstruction from Tongue and Lip Articulation with Transfer Learning from Text-to-Speech Synthesis

AAAI 2021technical

This paper presents TaLNet, a model for voice reconstruction with ultrasound tongue and optical lip videos as inputs. TaLNet is based on an encoder-decoder architecture. Separate encoders are dedicated to processing the tongue and lip data streams respectively. The decoder pre…

Cited by 18SourcePDFScholar
2020

A Tree-Structured Decoder for Image-to-Markup Generation

ICML 2020poster

Recent encoder-decoder approaches typically employ string decoders to convert images into serialized strings for image-to-markup. However, for tree-structured representational markup, string representations can hardly cope with the structural complexity. In this work, we first show via a set of toy…

Cited by 89SourcePDFScholar
2020

An Improved Deep Neural Network for Modeling Speaker Characteristics at Different Temporal Scales

ICASSP 2020accepted

This paper presents an improved deep embedding learning method based on a convolutional neural network (CNN) for text-independent speaker verification. Two improvements are proposed for x-vector embedding learning: (1) a multiscale convolution (MSCNN) is adopted in the frame-level layers to capture…

Cited by 0SourceScholar
2020

Attention-Based Gated Scaling Adaptive Acoustic Model for CTC-Based Speech Recognition

ICASSP 2020accepted

In this paper, we propose a novel adaptive technique that uses an attention-based gated scaling (AGS) scheme to improve deep feature learning for connectionist temporal classification (CTC) acoustic modeling. In AGS, the outputs of each hidden layer of the main network are scaled by an auxiliary gat…

Cited by 0SourceScholar
2018

Pseudo-Supervised Approach for Text Clustering Based on Consensus Analysis

ICASSP 2018accepted

In recent years, neural networks (NN) have achieved remarkable performance improvement in text classification due to their powerful ability to encode discriminative features by incorporating label information into model training. Inspired by the success of NN in text classification, we propose a pse…

Cited by 0SourceScholar