← Search

Tan Lee

24 accepted papers

2026

Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models

ICLR 2026poster

Large Audio Language Models (LALMs) represent an important frontier in multimodal AI, addressing diverse audio tasks. Recently, post-training of LALMs has received increasing attention due to significant performance improvements over foundation models. While single-stage post-training such as reinfo…

Cited by 0SourcecodeScholar
2025

PodAgent: A Comprehensive Framework for Podcast Generation

ACL 2025finding

Existing automatic audio generation methods struggle to generate podcast-like audio programs effectively. The key challenges lie in in-depth content generation, appropriate and expressive voice production. This paper proposed PodAgent, a comprehensive framework for creating audio programs. PodAgent…

2024

Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss

ICASSP 2024accepted

This research is about the creation of personalized synthetic voices for head and neck cancer survivors. It is focused particularly on tongue cancer patients whose speech might exhibit severe articulation impairment. Our goal is to restore normal articulation in the synthesized speech, while maximal…

Cited by 0SourceScholar
2024

Efficient Black-Box Speaker Verification Model Adaptation With Reprogramming And Backend Learning

ICASSP 2024accepted

The development of deep neural networks (DNN) has significantly enhanced the performance of speaker verification (SV) systems in recent years. However, a critical issue that persists when applying DNN-based SV systems in practical applications is domain mismatch. To mitigate the performance degradat…

Cited by 0SourceScholar
2024

Modeling Intrapersonal and Interpersonal Influences for Automatic Estimation of Therapist Empathy in Counseling Conversation

ICASSP 2024accepted

Counseling is usually conducted through spoken conversation between a therapist and a client. The empathy level of therapist is a key indicator of outcomes. Presuming that therapist’s empathy expression is shaped by their past behavior and their perception of the client’s behavior, we propose a mode…

Cited by 0SourceScholar
2023

An ASR-Free Fluency Scoring Approach with Self-Supervised Learning

ICASSP 2023accepted

A typical fluency scoring system generally relies on an automatic speech recognition (ASR) system to obtain time stamps in input speech for the subsequent calculation of fluency-related features or directly modeling speech fluency with an end-to-end approach. This paper describes a novel ASR-free ap…

Cited by 0SourceScholar
2023

Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification

ICASSP 2023accepted

Deep convolutional neural networks (CNNs) have been applied to extracting speaker embeddings with significant success in speaker verification. Incorporating the attention mechanism has shown to be effective in improving the model performance. This paper presents an efficient two-dimensional convolut…

Cited by 0SourceScholar
2023

Covariance Regularization for Probabilistic Linear Discriminant Analysis

ICASSP 2023accepted

Probabilistic linear discriminant analysis (PLDA) is commonly used in speaker verification systems to score the similarity of speaker embeddings. Recent studies improved the performance of PLDA in domain-matched conditions by diagonalizing its covariance. We suspect such a brutal pruning approach co…

Cited by 0SourceScholar
2023

Leveraging Phone-Level Linguistic-Acoustic Similarity For Utterance-Level Pronunciation Scoring

ICASSP 2023accepted

Recent studies on pronunciation scoring have explored the effect of introducing phone embeddings as reference pronunciation, but mostly in an implicit manner, i.e., addition or concatenation of reference phone embedding and actual pronunciation of the target phone as the phone-level pronunciation qu…

Cited by 0SourceScholar
2022

A Study on the Efficacy of Model Pre-Training In Developing Neural Text-to-Speech System

ICASSP 2022accepted

In the development of neural text-to-speech systems, model pre-training with a large amount of non-target speakers’ data is a common approach. However, in terms of ultimately achieved system performance for target speaker(s), the actual benefits of model pre-training are uncertain and unstable, depe…

Cited by 0SourceScholar
2020

Mixture Factorized Auto-Encoder for Unsupervised Hierarchical Deep Factorization of Speech Signal

ICASSP 2020accepted

Speech signal is constituted and contributed by various informative factors, such as linguistic content and speaker characteristic. There have been notable recent studies attempting to factorize speech signal into these individual factors without requiring any annotation. These studies typically ass…

Cited by 0SourceScholar
2020

Resting-State EEG-Based Biometrics with Signals Features Extracted by Multivariate Empirical Mode Decomposition

ICASSP 2020accepted

EEG-based biometrics has gained great attention in recent years due to its superiority over traditional biometrics in terms of its resistance to circumvention. While there are numerous choices of data acquisition protocol, the present study is carried out with the least demanding resting-state condi…

Cited by 0SourceScholar
2019

Adversarial Multi-task Deep Features and Unsupervised Back-end Adaptation for Language Recognition

ICASSP 2019accepted

This paper presents an investigation into speaker-invariant feature learning and domain adaptation for language recognition (LR) with short utterances. While following the conventional design of i-vector front-end and probabilistic linear discriminant analysis (PLDA) back-end, we propose to apply sp…

Cited by 0SourceScholar
2019

BLHUC: Bayesian Learning of Hidden Unit Contributions for Deep Neural Network Speaker Adaptation

ICASSP 2019accepted

Speaker adaptation techniques play a key role in reducing the mismatch between speech recognition systems and target users. In order to robustly learn speaker-dependent adaptation parameters, model based DNN adaptation techniques often require a significant amount of data. For example, in the common…

Cited by 0SourceScholar
2019

Combining Phone Posteriorgrams from Strong and Weak Recognizers for Automatic Speech Assessment of People with Aphasia

ICASSP 2019accepted

This paper presents an investigation on applying automatic speech recognition (ASR) to speech assessment of people with aphasia (PWA). A distinctive characteristic of PWA speech is paraphasia, which refers to frequent occurrence of phonemic errors, unintended words and non-verbal sounds. In view of…

Cited by 0SourceScholar
2019

Revisiting Hidden Markov Models for Speech Emotion Recognition

ICASSP 2019accepted

Hidden Markov models (HMMs) have a long tradition in automatic speech recognition (ASR) due to their capability of capturing temporal dynamic characteristics of speech. For emotion recognition from speech, three HMM based architectures are investigated and compared throughout the current paper, name…

Cited by 0SourceScholar
2018

Automatic Speech Assessment for Aphasic Patients Based on Syllable-Level Embedding and Supra-Segmental Duration Features

ICASSP 2018accepted

Aphasia is a type of acquired language impairment resulting from brain injury. Speech assessment is an important part of the comprehensive assessment process for aphasic patients. It is based on the acoustical and linguistic analysis of patients' speech elicited through pre-defined story-telling tas…

Cited by 0SourceScholar
2017

Polyphonic piano note transcription with non-negative matrix factorization of differential spectrogram

ICASSP 2017accepted

Automatic music transcription is usually approached by using a time-frequency (TF) representation such as the short-time Fourier transform (STFT) spectrogram or the constant-Q transform. In this paper, we propose a novel yet simple TF representation that capitalizes the effectiveness of spectral flu…

Cited by 0SourceScholar
2017

Shefce: A Cantonese-English bilingual speech corpus for pronunciation assessment

ICASSP 2017accepted

This paper introduces the development of ShefCE: a Cantonese-English bilingual speech corpus from L2 English speakers in Hong Kong. Bilingual parallel recording materials were chosen from TED online lectures. Script selection were carried out according to bilingual consistency (evaluated using a mac…

Cited by 0SourceScholar
2016

Automatic speech recognition for acoustical analysis and assessment of cantonese pathological voice and speech

ICASSP 2016accepted

This paper describes the application of state-of-the-art automatic speech recognition (ASR) systems to objective assessment of voice and speech disorders. Acoustical analysis of speech has long been considered a promising approach to non-invasive and objective assessment of people. In the past the t…

Cited by 0SourceScholar