← Search

Tamás Grósz

4 accepted papers

2024

Collecting Linguistic Resources for Assessing Children’s Pronunciation of Nordic Languages

COLING 2024main

This paper reports on the experience collecting a number of corpora of Nordic languages spoken by children. The aim of the data collection is providing annotated data to develop and evaluate computer assisted pronunciation assessment systems both for non-native children learning a Nordic language (L…

Cited by 2SourcePDFScholar
2024

Investigating the Clusters Discovered By Pre-Trained AV-HuBERT

ICASSP 2024accepted

Self-supervised models, such as HuBERT and its audio-visual version AV-HuBERT, have demonstrated excellent performance on various tasks. The main factor for their success is the pre-training procedure, which requires only raw data without human transcription. During the self-supervised pre-training…

Cited by 0SourceScholar
2018

F0 Estimation for DNN-Based Ultrasound Silent Speech Interfaces

ICASSP 2018accepted

State-of-the-art silent speech interface systems apply vocoders to generate the speech signal directly from articulatory data. Most of these approaches concentrate on estimating just the spectral features of the vocoder, and use the original F0, a constant F0 or white noise as excitation. This solut…

Cited by 0SourceScholar
2015

Building context-dependent DNN acoustic models using Kullback-Leibler divergence-based state tying

ICASSP 2015accepted

Deep neural network (DNN) based speech recognizers have recently replaced Gaussian mixture (GMM) based systems as the state-of-the-art. HMM/DNN systems have kept many refinements of the HMM/GMM framework, even though some of these may be suboptimal for them. One such example is the creation of conte…

Cited by 0SourceScholar