← Search

Kentaro Mitsui

3 accepted papers

2024

Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition

ACL 2024findings

Advances in machine learning have made it possible to perform various text and speech processing tasks, such as automatic speech recognition (ASR), in an end-to-end (E2E) manner. E2E approaches utilizing pre-trained models are gaining attention for conserving training data and resources. However, mo…

2024

PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems

EMNLP 2024finding

Multimodal language models that process both text and speech have a potential for applications in spoken dialogue systems. However, current models face two major challenges in response generation latency: (1) generating a spoken response requires the prior generation of a written response, and (2) s…

2024

Release of Pre-Trained Models for the Japanese Language

COLING 2024main

AI democratization aims to create a world in which the average person can utilize AI techniques. To achieve this goal, numerous research institutes have attempted to make their results accessible to the public. In particular, large pre-trained models trained on large-scale data have shown unpreceden…

Cited by 18SourcePDFScholar