← Search

Anton Ragni

12 accepted papers

2024

MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

ICLR 2024poster

Self-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech. Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored.…

2024

Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users Using Intermediate ASR Features and Human Memory Models

ICASSP 2024accepted

Neural networks have been successfully used for non-intrusive speech intelligibility prediction. Recently, the use of feature representations sourced from intermediate layers of pre-trained self-supervised and weakly-supervised models has been found to be particularly useful for this task. This work…

Cited by 0SourceScholar
2020

Confidence Estimation for Black Box Automatic Speech Recognition Systems Using Lattice Recurrent Neural Networks

ICASSP 2020accepted

Recently, there has been growth in providers of speech transcription services enabling others to leverage technology they would not normally be able to use. As a result, speech-enabled solutions have become commonplace. Their success critically relies on the quality, accuracy, and reliability of the…

Cited by 0SourceScholar
2019

Bi-directional Lattice Recurrent Neural Networks for Confidence Estimation

ICASSP 2019accepted

The standard approach to mitigate errors made by an automatic speech recognition system is to use confidence scores associated with each predicted word. In the simplest case, these scores are word posterior probabilities whilst more complex schemes utilise bi-directional recurrent neural network (Bi…

Cited by 0SourceScholar
2018

Phonetic and Graphemic Systems for Multi-Genre Broadcast Transcription

ICASSP 2018accepted

State-of-the-art English automatic speech recognition systems typically use phonetic rather than graphemic lexicons. Graphemic systems are known to perform less well for English as the mapping from the written form to the spoken form is complicated. However, in recent years the representational powe…

Cited by 0SourceScholar
2017

Morph-to-word transduction for accurate and efficient automatic speech recognition and keyword search

ICASSP 2017accepted

Word units are a popular choice in statistical language modelling. For inflective and agglutinative languages this choice may result in a high out of vocabulary rate. Subword units, such as morphs, provide an interesting alternative to words. These units can be derived in an unsupervised fashion and…

Cited by 0SourceScholar
2017

Recurrent neural network language models for keyword search

ICASSP 2017accepted

Recurrent neural network language models (RNNLMs) have becoming increasingly popular in many applications such as automatic speech recognition (ASR). Significant performance improvements in both perplexity and word error rate over standard n-gram LMs have been widely reported on ASR tasks. In contra…

Cited by 0SourceScholar
2017

Stimulated training for automatic speech recognition and keyword search in limited resource conditions

ICASSP 2017accepted

Training neural network acoustic models on limited quantities of data is a challenging task. A number of techniques have been proposed to improve generalisation. This paper investigates one such technique called stimulated training. It enables standard criteria such as cross-entropy to enforce spati…

Cited by 0SourceScholar