← Search

Jari Kolehmainen

4 accepted papers

2025

Speech Recognition Rescoring with Large Speech-Text Foundation Models

ICASSP 2025accepted

Large language models (LLM) have demonstrated the ability to understand human language by leveraging large amount of text data. Automatic speech recognition (ASR) systems are often limited by available transcribed speech data and benefit from a second pass rescoring using LLM. Recently multi-modal l…

Cited by 0SourceScholar
2024

Multi-Modal Retrieval For Large Language Model Based Speech Recognition

ACL 2024findings

Retrieval is a widely adopted approach for improving language models leveraging external information. As the field moves towards multi-modal large language models, it is important to extend the pure text based methods to incorporate other modalities in retrieval as well for applications across the w…

2024

Towards ASR Robust Spoken Language Understanding Through in-Context Learning with Word Confusion Networks

ICASSP 2024accepted

In the realm of spoken language understanding (SLU). numerous natural language understanding (NLU) methodologies have been adapted by supplying large language models (LLMs) with transcribed speech instead of conventional written text. In real-world scenarios, prior to input into an LLM. an automated…

Cited by 0SourceScholar
2022

RescoreBERT: Discriminative Speech Recognition Rescoring With Bert

ICASSP 2022accepted

Second-pass rescoring is an important component in automatic speech recognition (ASR) systems that is used to improve the outputs from a first-pass decoder by implementing a lattice rescoring or n-best re-ranking. While pretraining with a masked language model (MLM) objective has received great succ…

Cited by 0SourceScholar