← Search

Jiahong Yuan

12 accepted papers

2025

Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models

ICASSP 2025accepted

Utilizing Self-Supervised Learning (SSL) models for Speech Emotion Recognition (SER) has proven effective, yet limited research has explored cross-lingual scenarios. This study presents a comparative analysis between human performance and SSL models, beginning with a layer-wise analysis and an explo…

Cited by 0SourceScholar
2025

The USTC System for EEG-Music Emotion Recognition Challenge

ICASSP 2025accepted

This paper presents the Neural Harmony team’s submission to Task 1 (Person Identification) of the ICASSP 2025 EEG-Music Emotion Recognition Challenge, which aims to identify the subject from a given EEG segment. To enhance performance, we propose a novel architecture incorporating the Multiscale Con…

Cited by 0SourceScholar
2025

Transformer-based Speech Model Learns Well as Infants and Encodes Abstractions through Exemplars in the Poverty of the Stimulus Environment

COLING 2025main

Infants are capable of learning language, predominantly through speech and associations, in impoverished environments—a phenomenon known as the Poverty of the Stimulus (POS). Is this ability uniquely human, as an innate linguistic predisposition, or can it be empirically learned through potential li…

2024

Automated Tone Transcription and Clustering with Tone2Vec

EMNLP 2024finding

Lexical tones play a crucial role in Sino-Tibetan languages. However, current phonetic fieldwork relies on manual effort, resulting in substantial time and financial costs. This is especially challenging for the numerous endangered languages that are rapidly disappearing, often compounded by limited…

2022

Text2video: Text-Driven Talking-Head Video Synthesis with Personalized Phoneme - Pose Dictionary

ICASSP 2022accepted

With the advance of deep learning technology, automatic video generation from audio or text has become an emerging and promising research topic. In this paper, we present a novel approach to synthesize video from the text. The method builds a phoneme-pose dictionary and trains a generative adversari…

Cited by 0SourceScholar
2022

W-CTC: a Connectionist Temporal Classification Loss with Wild Cards

ICLR 2022poster

Connectionist Temporal Classification (CTC) loss is commonly used in sequence learning applications. For example, in Automatic Speech Recognition (ASR) task, the training data consists of pairs of audio (input sequence) and text (output label),without temporal alignment information. Standard CTC com…

Cited by 10SourcePDFScholar
2021

On Attention Redundancy: A Comprehensive Study

NAACL 2021long

Multi-layer multi-head self-attention mechanism is widely applied in modern neural language models. Attention redundancy has been observed among attention heads but has not been deeply studied in the literature. Using BERT-base model as an example, this paper provides a comprehensive study on attent…

2021

Pause-Encoded Language Models for Recognition of Alzheimer's Disease and Emotion

ICASSP 2021accepted

We propose enhancing Transformer language models (BERT, RoBERTa) to take advantage of pauses. Pauses play an important role in speech. In previous work we developed a method to encode pauses in transcripts for recognition of Alzheimer's disease. In this study, we extend this idea to language models.…

Cited by 0SourceScholar
2021

Speaking Rate and Tonal Realization in Mandarin Chinese: What Can We Learn From Large Speech Corpora?

ICASSP 2021accepted

Two Mandarin speech corpora were used to investigate tonal realization in terms of duration and pitch. The data consist of nearly 1000 hours of speech from more than 1600 speakers. The two corpora, both developed for ASR, differ in speaking rate by approximately 25%. This provides an opportunity to…

Cited by 0SourceScholar