← Search

Jinglin Liu

26 accepted papers

2024

AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

AAAI 2024technical

Large language models (LLMs) have exhibited remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. Despite the recent success, current LLMs are not capable of processing complex audio information or conducting spoken conversations (lik…

2024

Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis

ICLR 2024poster

Zero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms of zero-shot TTS still face challenges in the following aspe…

2024

MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes

NeurIPS 2024poster

Talking face generation (TFG) aims to animate a target identity's face to create realistic talking videos. Personalized TFG is a variant that emphasizes the perceptual identity similarity of the synthesized result (from the perspective of appearance and talking style). While previous works typically…

2024

Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

ICLR 2024spotlight

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the goals of accurate 3D avatar reconstruction and stable talkin…

2023

AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation

ACL 2023long

Direct speech-to-speech translation (S2ST) aims to convert speech from one language into another, and has demonstrated significant progress to date. Despite the recent success, current S2ST models still suffer from distinct degradation in noisy environments and fail to translate visual speech (i.e.,…

2023

AlignSTS: Speech-to-Singing Conversion via Cross-Modal Alignment

ACL 2023findings

The speech-to-singing (STS) voice conversion task aims to generate singing samples corresponding to speech recordings while facing a major challenge: the alignment between the target (singing) pitch contour and the source (speech) content is difficult to learn in a text-free situation. This paper pr…

2023

CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-Training

ACL 2023long

Improving text representation has attracted much attention to achieve expressive text-to-speech (TTS). However, existing works only implicitly learn the prosody with masked token reconstruction tasks, which leads to low training efficiency and difficulty in prosody modeling. We propose CLAPSpeech, a…

2023

DopplerBAS: Binaural Audio Synthesis Addressing Doppler Effect

ACL 2023findings

Recently, binaural audio synthesis (BAS) has emerged as a promising research field for its applications in augmented and virtual realities. Binaural audio helps ususers orient themselves and establish immersion by providing the brain with interaural time differences reflecting spatial information. H…

Cited by 1SourcePDFScholar
2023

FastDiff 2: Revisiting and Incorporating GANs and Diffusion Models in High-Fidelity Speech Synthesis

ACL 2023findings

Generative adversarial networks (GANs) and denoising diffusion probabilistic models (DDPMs) have recently achieved impressive performances in image and audio synthesis. After revisiting their success in conditional speech synthesis, we find that 1) GANs sacrifice sample diversity for quality and spe…

2023

GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis

ICLR 2023poster

Generating photo-realistic video portraits with arbitrary speech audio is a crucial problem in film-making and virtual reality. Recently, several works explore the usage of neural radiance field (NeRF) in this task to improve 3D realness and image fidelity. However, the generalizability of previous…

2023

MUG: A General Meeting Understanding and Generation Benchmark

ICASSP 2023accepted

Listening to long video/audio recordings from video conferencing and online courses for acquiring information is extremely inefficient. Even after ASR systems transcribe recordings into long-form spoken language documents, reading ASR transcripts only partly speeds up seeking information. It has bee…

Cited by 0SourceScholar
2023

Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

ICML 2023poster

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high-quality text-audio pairs, and the complexity of modeling long continuous audio…

2023

Overview of the ICASSP 2023 General Meeting Understanding and Generation Challenge (MUG)

ICASSP 2023accepted

ICASSP2023 General Meeting Understanding and Generation Challenge (MUG) focuses on prompting a wide range of spoken language processing (SLP) research on meeting transcripts, as SLP applications are critical to improve users’ efficiency in grasping important information in meetings. MUG includes fiv…

Cited by 0SourceScholar
2023

RMSSinger: Realistic-Music-Score based Singing Voice Synthesis

ACL 2023findings

We are interested in a challenging task, Realistic-Music-Score based Singing Voice Synthesis (RMS-SVS). RMS-SVS aims to generate high-quality singing voices given realistic music scores with different note types (grace, slur, rest, etc.). Though significant progress has been achieved, recent singing…

2023

TranSpeech: Speech-to-Speech Translation With Bilateral Perturbation

ICLR 2023poster

Direct speech-to-speech translation (S2ST) with discrete units leverages recent progress in speech representation learning. Specifically, a sequence of discrete representations derived in a self-supervised manner are predicted from the model and passed to a vocoder for speech reconstruction, while s…

2023

VarietySound: Timbre-Controllable Video to Sound Generation Via Unsupervised Information Disentanglement

ICASSP 2023accepted

Video-to-sound generation aims to generate realistic and natural sound given a video input. However, previous video-to-sound generation methods can only generate a random or average timbre without any controls of the generated sound timbre, leading to the problem that people cannot obtain the desire…

Cited by 0SourceScholar
2022

Dict-TTS: Learning to Pronounce with Prior Dictionary Knowledge for Text-to-Speech

NeurIPS 2022accept

Polyphone disambiguation aims to capture accurate pronunciation knowledge from natural text sequences for reliable Text-to-speech (TTS) systems. However, previous approaches require substantial annotated training data and additional efforts from language experts, making it difficult to extend high-q…

2022

DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism

AAAI 2022technical

Singing voice synthesis (SVS) systems are built to synthesize high-quality and expressive singing voice, in which the acoustic model generates the acoustic features (e.g., mel-spectrogram) given a music score. Previous singing acoustic models adopt a simple loss (e.g., L1 and L2) or generative adver…

2022

Flow-Based Unconstrained Lip to Speech Generation

AAAI 2022technical

Unconstrained lip-to-speech aims to generate corresponding speeches based on silent facial videos with no restriction to head pose or vocabulary. It is desirable to generate intelligible and natural speech with a fast speed in unconstrained settings. Currently, to handle the more complicated scena…

2022

GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-Speech

NeurIPS 2022accept

Style transfer for out-of-domain (OOD) speech synthesis aims to generate speech samples with unseen style (e.g., speaker identity, emotion, and prosody) derived from an acoustic reference, while facing the following challenges: 1) The highly dynamic style features in expressive voice are difficult t…

2022

Learning the Beauty in Songs: Neural Singing Voice Beautifier

ACL 2022long

We are interested in a novel task, singing voice beautification (SVB). Given the singing voice of an amateur singer, SVB aims to improve the intonation and vocal tone of the voice, while keeping the content and vocal timbre. Current automatic pitch correction techniques are immature, and most of the…

2022

M4Singer: A Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing Corpus

NeurIPS 2022accept

The lack of publicly available high-quality and accurately labeled datasets has long been a major bottleneck for singing voice synthesis (SVS). To tackle this problem, we present M4Singer, a free-to-use Multi-style, Multi-singer Mandarin singing collection with elaborately annotated Musical scores a…

Cited by 89SourcePDFScholar
2022

Parallel and High-Fidelity Text-to-Lip Generation

AAAI 2022technical

As a key component of talking face generation, lip movements generation determines the naturalness and coherence of the generated talking face video. Prior literature mainly focuses on speech-to-lip generation while there is a paucity in text-to-lip (T2L) generation. T2L is a challenging task and ex…

2021

Denoispeech: Denoising Text to Speech with Frame-Level Noise Modeling

ICASSP 2021accepted

While neural-based text to speech (TTS) models can synthesize natural and intelligible voice, they usually require high-quality speech data, which is costly to collect. In many scenarios, only noisy speech of a target speaker is available, which presents challenges for TTS model training for this sp…

Cited by 0SourceScholar
2020

Task-Level Curriculum Learning for Non-Autoregressive Neural Machine Translation

IJCAI 2020poster

Non-autoregressive translation (NAT) achieves faster inference speed but at the cost of worse accuracy compared with autoregressive translation (AT). Since AT and NAT can share model structure and AT is an easier task than NAT due to the explicit dependency on previous target-side tokens, a natural…