← Search

Ryandhimas E. Zezario

4 accepted papers

2026

FEW-SHOT AND PSEUDO-LABEL GUIDED SPEECH QUALITY EVALUATION WITH LARGE LANGUAGE MODELS

ICASSP 2026oral

In this paper, we introduce GatherMOS, a novel framework that leverages large language models (LLM) as meta-evaluators to aggregate diverse signals into quality predictions. GatherMOS integrates lightweight acoustic descriptors with pseudo-labels from DNSMOS and VQScore, enabling the LLM to reason o…

Cited by 0SourcePDFScholar
2025

A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models

ICASSP 2025accepted

This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which uses Whisper as an audio-to-text module and evaluates the text’s naturalness via ta…

Cited by 0SourceScholar
2024

Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment Model

ICASSP 2024accepted

This study proposes a multi-task pseudo-label learning (MPL)-based non-intrusive speech quality assessment model called MTQ-Net. MPL consists of two stages: obtaining pseudo-label scores from a pretrained model and performing multitask learning. The 3QUEST metrics, namely Speech-MOS (S-MOS), Noise-M…

Cited by 0SourceScholar
2020

Self-Supervised Denoising Autoencoder with Linear Regression Decoder for Speech Enhancement

ICASSP 2020accepted

Nonlinear spectral mapping-based models based on supervised learning have successfully applied for speech enhancement. However, as supervised learning approaches, a large amount of labelled data (noisy-clean speech pairs) should be provided to train those models. In addition, their performances for…

Cited by 0SourceScholar