← Search

Hsin-Min Wang

27 accepted papers

2026

FEW-SHOT AND PSEUDO-LABEL GUIDED SPEECH QUALITY EVALUATION WITH LARGE LANGUAGE MODELS

ICASSP 2026oral

In this paper, we introduce GatherMOS, a novel framework that leverages large language models (LLM) as meta-evaluators to aggregate diverse signals into quality predictions. GatherMOS integrates lightweight acoustic descriptors with pseudo-labels from DNSMOS and VQScore, enabling the LLM to reason o…

Cited by 0SourcePDFScholar
2025

A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models

ICASSP 2025accepted

This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which uses Whisper as an audio-to-text module and evaluates the text’s naturalness via ta…

Cited by 0SourceScholar
2025

Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition

ICASSP 2025accepted

While pre-trained automatic speech recognition (ASR) systems demonstrate impressive performance on matched domains, their performance often degrades when confronted with channel mismatch stemming from unseen recording environments and conditions. To mitigate this issue, we propose a novel channel-aw…

Cited by 0SourceScholar
2025

Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement

ICASSP 2025accepted

In multichannel speech enhancement, effectively capturing spatial and spectral information across different microphones is crucial for noise reduction. Traditional methods, such as CNN or LSTM, attempt to model the temporal dynamics of full-band and sub-band spectral and spatial features. However, t…

Cited by 0SourceScholar
2025

MSECG: Incorporating Mamba for Robust and Efficient ECG Super-Resolution

ICASSP 2025accepted

Electrocardiogram (ECG) signals play a crucial role in diagnosing cardiovascular diseases. To reduce power consumption in wearable or portable devices used for long-term ECG monitoring, super-resolution (SR) techniques have been developed, enabling these devices to collect and transmit signals at a…

Cited by 0SourceScholar
2024

Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment Model

ICASSP 2024accepted

This study proposes a multi-task pseudo-label learning (MPL)-based non-intrusive speech quality assessment model called MTQ-Net. MPL consists of two stages: obtaining pseudo-label scores from a pretrained model and performing multitask learning. The 3QUEST metrics, namely Speech-MOS (S-MOS), Noise-M…

Cited by 0SourceScholar
2023

D4AM: A General Denoising Framework for Downstream Acoustic Models

ICLR 2023poster

The performance of acoustic models degrades notably in noisy environments. Speech enhancement (SE) can be used as a front-end strategy to aid automatic speech recognition (ASR) systems. However, existing training objectives of SE methods are not fully effective at integrating speech-text and noise-c…

2022

Partially Fake Audio Detection by Self-Attention-Based Fake Span Discovery

ICASSP 2022accepted

The past few years have witnessed the significant advances of speech synthesis and voice conversion technologies. However, such technologies can undermine the robustness of broadly implemented biometric identification models and can be harnessed by in-the-wild attackers for illegal uses. The ASVspoo…

Cited by 0SourceScholar
2021

Melody Harmonization Using Orderless Nade, Chord Balancing, and Blocked Gibbs Sampling

ICASSP 2021accepted

Coherence and interestingness are two criteria for evaluating the performance of melody harmonization, which aims to generate a chord progression from a symbolic melody. In this study, we apply the concept of orderless NADE, which takes the melody and its partially masked chord sequence as the input…

Cited by 0SourceScholar
2021

Sequence to General Tree: Knowledge-Guided Geometry Word Problem Solving

ACL 2021short

With the recent advancements in deep learning, neural solvers have gained promising results in solving math word problems. However, these SOTA solvers only generate binary expression trees that contain basic arithmetic operators and do not explicitly use the math formulas. As a result, the expressio…

2021

Speech Recognition by Simply Fine-Tuning Bert

ICASSP 2021accepted

We propose a simple method for automatic speech recognition (ASR) by fine-tuning BERT, which is a language model (LM) trained on large-scale unlabeled text data and can generate rich contextual representations. Our assumption is that given a history context sequence, a powerful LM can narrow the ran…

Cited by 0SourceScholar
2020

Combining Deep Embeddings of Acoustic and Articulatory Features for Speaker Identification

ICASSP 2020accepted

In this study, deep embedding of acoustic and articulatory features are combined for speaker identification. First, a convolutional neural network (CNN)-based universal background model (UBM) is constructed to generate acoustic feature (AC) embedding. In addition, as the articulatory features (AFs)…

Cited by 0SourceScholar
2020

Self-Supervised Denoising Autoencoder with Linear Regression Decoder for Speech Enhancement

ICASSP 2020accepted

Nonlinear spectral mapping-based models based on supervised learning have successfully applied for speech enhancement. However, as supervised learning approaches, a large amount of labelled data (noisy-clean speech pairs) should be provided to train those models. In addition, their performances for…

Cited by 0SourceScholar
2020

Statistics Pooling Time Delay Neural Network Based on X-Vector for Speaker Verification

ICASSP 2020accepted

This paper aims to improve speaker embedding representation based on x-vector for extracting more detailed information for speaker verification. We propose a statistics pooling time delay neural network (TDNN), in which the TDNN structure integrates statistics pooling for each layer, to consider the…

Cited by 0SourceScholar
2019

Reinforcement Learning Based Speech Enhancement for Robust Speech Recognition

ICASSP 2019accepted

Conventional deep neural network (DNN)-based speech enhancement (SE) approaches aim to minimize the mean square error (MSE) between enhanced speech and clean reference. The MSE-optimized model may not directly improve the performance of an automatic speech recognition (ASR) system. If the target is…

Cited by 32SourceScholar
2018

Essence Vector-Based Query Modeling for Spoken Document Retrieval

ICASSP 2018accepted

Spoken document retrieval (SDR) has become a prominently required application since unprecedented volumes of multimedia data along with speech have become available in our daily life. As far as we are aware, there has been relatively less work in launching unsupervised paragraph embedding methods an…

Cited by 0SourceScholar
2017

A locality-preserving essence vector modeling framework for spoken document retrieval

ICASSP 2017accepted

Because unprecedented volumes of multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research area in the past decades. Recently, representation learning has emerged as an active research topic in many machi…

Cited by 0SourceScholar
2017

A locally linear embbeding based postfiltering approach for speech enhancement

ICASSP 2017accepted

This paper presents a novel postfiltering approach based on the locally linear embedding (LLE) algorithm for speech enchantment (SE). The aim of the proposed LLE-based postfiltering approach is to further remove the residual noise components from the SE-processed speech signals through a spectral co…

Cited by 0SourceScholar
2017

Deep-net fusion to classify shots in concert videos

ICASSP 2017accepted

Varying types of shots is a fundamental element in the language of film, commonly used by a visual storytelling director to convey the emotion, ideas, and art. To classify such types of shots from images, we present a new framework that facilitates the intriguing task by addressing two key issues. W…

Cited by 0SourceScholar
2017

Discriminative autoencoders for speaker verification

ICASSP 2017accepted

This paper presents a learning and scoring framework based on neural networks for speaker verification. The framework employs an autoencoder as its primary structure while three factors are jointly considered in the objective function for speaker discrimination. The first one, relating to the sample…

Cited by 0SourceScholar
2017

Leveraging manifold learning for extractive broadcast news summarization

ICASSP 2017accepted

Extractive speech summarization is intended to produce a condensed version of the original spoken document by selecting a few salient sentences from the document and concatenate them together to form a summary. In this paper, we study a novel use of manifold learning techniques for extractive speech…

Cited by 0SourceScholar
2016

DEMV-matchmaker: Emotional temporal course representation and deep similarity matching for automatic music video generation

ICASSP 2016accepted

This paper presents a deep similarity matching-based emotion-oriented music video (MV) generation system, called DEMV-matchmaker, which utilizes an emotion-oriented deep similarity matching (EDSM) metric as a bridge to connect music and video. Specifically, we adopt an emotional temporal course mode…

Cited by 12SourceScholar
2016

Improved spoken document summarization with coverage modeling techniques

ICASSP 2016accepted

Extractive summarization aims at selecting a set of indicative sentences from a source document as a summary that can express the major theme of the document. A general consensus on extractive summarization is that both relevance and coverage are critical issues to address. The existing methods desi…

Cited by 0SourceScholar
2015

A histogram density modeling approach to music emotion recognition

ICASSP 2015accepted

Music emotion recognition is concerned with developing predictive models that comprehend the affective content of musical signals. Recently, a growing number of attempts has been made to model the music emotion as a probability distribution in the valence-arousal (VA) space to better account for the…

Cited by 0SourceScholar