← Search

Mikko Kurimo

8 accepted papers

2024

Collecting Linguistic Resources for Assessing Children’s Pronunciation of Nordic Languages

COLING 2024main

This paper reports on the experience collecting a number of corpora of Nordic languages spoken by children. The aim of the data collection is providing annotated data to develop and evaluate computer assisted pronunciation assessment systems both for non-native children learning a Nordic language (L…

Cited by 2SourcePDFScholar
2024

Investigating the Clusters Discovered By Pre-Trained AV-HuBERT

ICASSP 2024accepted

Self-supervised models, such as HuBERT and its audio-visual version AV-HuBERT, have demonstrated excellent performance on various tasks. The main factor for their success is the pre-training procedure, which requires only raw data without human transcription. During the self-supervised pre-training…

Cited by 0SourceScholar
2022

When to Laugh and How Hard? A Multimodal Approach to Detecting Humor and Its Intensity

COLING 2022main

Prerecorded laughter accompanying dialog in comedy TV shows encourages the audience to laugh by clearly marking humorous moments in the show. We present an approach for automatically detecting humor in the Friends TV show using multimodal data. Our model is capable of recognizing whether an utteranc…

2021

Vowel Non-Vowel Based Spectral Warping and Time Scale Modification for Improvement in Children's ASR

ICASSP 2021accepted

Acoustic differences between children’s and adults’ speech causes the degradation in the automatic speech recognition system performance when system trained on adults’ speech and tested on children’s speech. The key acoustic mismatch factors are formant, speaking rate, and pitch. In this paper, we p…

Cited by 0SourceScholar
2020

Speaker-Aware Training of Attention-Based End-to-End Speech Recognition Using Neural Speaker Embeddings

ICASSP 2020accepted

In speaker-aware training, a speaker embedding is appended to DNN input features. This allows the DNN to effectively learn representations, which are robust to speaker variability.We apply speaker-aware training to attention-based end-to-end speech recognition. We show that it can improve over a pur…

Cited by 0SourceScholar
2020

Study of Formant Modification for Children ASR

ICASSP 2020accepted

The performance of automatic speech recognition systems for children’s speech is known to suffer from the large variation and mismatch in the acoustic and linguistic attributes between children’s and adults’ speech. One of the various identified sources of mismatch is the difference in formant frequ…

Cited by 0SourceScholar
2017

LDA-based context dependent recurrent neural network language model using document-based topic distribution of words

ICASSP 2017accepted

Adding context information into recurrent neural network language models (RNNLMs) have been investigated recently to improve the effectiveness of learning RNNLM. Conventionally, a fast approximate topic representation for a block of words was proposed by using corpus-based topic distribution of word…

Cited by 0SourceScholar
2015

Designing multichannel source separation based on single-channel source separation

ICASSP 2015accepted

In this paper, an extension of independent vector analysis (IVA), model-based IVA, is proposed for multichannel source separation. For obtaining better source models, we introduce a single-channel source separation method, and utilize the outputs as source variances in time-frequency-variant Gaussia…

Cited by 0SourceScholar