← Search

Changhe Song

6 accepted papers

2025

Binary Representation Learning for Discriminative Acoustic Unit Discovery

ICASSP 2025accepted

Acoustic Unit Discovery (AUD) aims to obtain phoneme-like units that preserve linguistically significant information while removing paralinguistic details. Although Contrastive Predictive Coding (CPC) has emerged as a leading self-supervised representation learning method for this task, CPC-based me…

Cited by 0SourceScholar
2025

Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States

ACL 2025long

As Large Language Models (LLMs) increasingly participate in human-AI interactions, evaluating their Theory of Mind (ToM) capabilities - particularly their ability to track dynamic mental states - becomes crucial. While existing benchmarks assess basic ToM abilities, they predominantly focus on stati…

2022

A Character-Level Span-Based Model for Mandarin Prosodic Structure Prediction

ICASSP 2022accepted

The accuracy of prosodic structure prediction is crucial to the naturalness of synthesized speech in Mandarin text-to-speech system, but now is limited by widely-used sequence-to-sequence framework and error accumulation from previous word segmentation results. In this paper, we propose a span-based…

Cited by 0SourceScholar
2022

An End-to-End Chinese Text Normalization Model Based on Rule-Guided Flat-Lattice Transformer

ICASSP 2022accepted

Text normalization, defined as a procedure transforming nonstandard words to spoken-form words, is crucial to the intelligibility of synthesized speech in text-to-speech system. Rule-based methods without considering context can not eliminate ambiguation, whereas sequence-to-sequence neural network…

Cited by 0SourceScholar
2022

Disentangling Content and Fine-Grained Prosody Information Via Hybrid ASR Bottleneck Features for Voice Conversion

ICASSP 2022accepted

Non-parallel data voice conversion (VC) have achieved considerable breakthroughs recently through introducing bottleneck features (BNFs) extracted by the automatic speech recognition(ASR) model. However, selection of BNFs have a significant impact on VC result. For example, when extracting BNFs from…

Cited by 0SourceScholar
2021

Syntactic Representation Learning For Neural Network Based TTS with Syntactic Parse Tree Traversal

ICASSP 2021accepted

Syntactic structure of a sentence text is correlated with the prosodic structure of the speech that is crucial for improving the prosody and naturalness of a text-to-speech (TTS) system. Nowadays TTS systems usually try to incorporate syntactic structure information with manually designed features b…

Cited by 0SourceScholar