← Search

Yibin Zheng

9 accepted papers

2026

Simulating Human-Like Counseling: A Path- and Scenario-Guided Framework for Psychological Support Dialogue

AAAI 2026technical

The growing demand for psychological support underscores the lack of high-quality counseling dialogue datasets, particularly in non-English contexts. We propose PGSim, a Path-Guided Simulation framework that mirrors real counseling processes—symptom description, problem identification, cause analysi

Cited by 0SourcePDFScholar
2025

QuASAR: A Question-Driven Structure-Aware Approach for Table-to-Text Generation

ACL 2025long

Table-to-text generation aims to automatically produce natural language descriptions from structured or semi-structured tabular data. Unlike traditional text generation tasks, it requires models to accurately understand and represent table structures. Existing approaches typically process tables by…

2023

WeSinger 2: Fully Parallel Singing Voice Synthesis via Multi-Singer Conditional Adversarial Training

ICASSP 2023accepted

This paper aims to introduce a robust singing voice synthesis (SVS) system to produce very natural and realistic singing voices efficiently by leveraging the adversarial training strategy. On one hand, we designed simple but generic random area conditional discriminators to help supervise the acoust…

Cited by 0SourceScholar
2022

Zero-Shot Cross-Lingual Transfer Using Multi-Stream Encoder and Efficient Speaker Representation

ICASSP 2022accepted

We propose a novel method for zero-shot cross-lingual TTS task by using multi-stream text encoder and efficient speaker representation. Specifically, a unified multi-stream text encoder that takes both advantages of Transformer and CBHG is proposed to retain multiple hypotheses about input represent…

Cited by 0SourceScholar
2021

Investigation of Fast and Efficient Methods for Multi-Speaker Modeling and Speaker Adaptation

ICASSP 2021accepted

In this paper, we propose a novel method for fast and efficient few-shot TTS task, which is able to disentangle linguistic and speaker representations. Specifically, an adversarial training strategy is firstly employed to wipe out speaker information from the linguistic representations. Then the spe…

Cited by 0SourceScholar
2020

An Improved Frame-Unit-Selection Based Voice Conversion System Without Parallel Training Data

ICASSP 2020accepted

A frame-unit-selection based voice conversion system proposed earlier by us is revisited here to enhance its performance in both speech naturalness and speaker similarity. Speaker independent, bilingual (Mandarin Chinese and American English) deep neural net (DNN) acoustic model’s output, frame-leve…

Cited by 1SourceScholar
2020

Improving End-to-End Speech Synthesis with Local Recurrent Neural Network Enhanced Transformer

ICASSP 2020accepted

Although Transformer based neural end-to-end TTS model has demonstrated extreme effectiveness in capturing long-term dependencies and achieved state-of-the-art performance, it still suffers from two problems. 1) limited ability to model sequential and local structures in sequences; 2) heavily rely o…

Cited by 0SourceScholar
2019

Phoneme Dependent Speaker Embedding and Model Factorization for Multi-speaker Speech Synthesis and Adaptation

ICASSP 2019accepted

This paper presents an architecture to perform speaker adaption in long short-term memory (LSTM) based Mandarin statistical parametric speech synthesis system. Compared with the conventional methods that focused on using fixed global speaker representations in utterance level for speaker recognition…

Cited by 0SourceScholar
2017

A novel pitch extraction based on jointly trained deep BLSTM Recurrent Neural Networks with bottleneck features

ICASSP 2017accepted

Pitch is an important characteristic of speech and is useful for many applications. However, it is still challenging to estimate pitch in strong noise. In this paper, we propose a joint training approach to determinate pitch. First, a Bidirectional Long Short-Term Memory Recurrent Neural Networks (B…

Cited by 0SourceScholar