← Search

Kei Sawada

7 accepted papers

2024

Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition

ACL 2024findings

Advances in machine learning have made it possible to perform various text and speech processing tasks, such as automatic speech recognition (ASR), in an end-to-end (E2E) manner. E2E approaches utilizing pre-trained models are gaining attention for conserving training data and resources. However, mo…

2024

PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems

EMNLP 2024finding

Multimodal language models that process both text and speech have a potential for applications in spoken dialogue systems. However, current models face two major challenges in response generation latency: (1) generating a spoken response requires the prior generation of a written response, and (2) s…

2024

Release of Pre-Trained Models for the Japanese Language

COLING 2024main

AI democratization aims to create a world in which the average person can utilize AI techniques. To achieve this goal, numerous research institutes have attempted to make their results accessible to the public. In particular, large pre-trained models trained on large-scale data have shown unpreceden…

Cited by 18SourcePDFScholar
2023

Focused Prefix Tuning for Controllable Text Generation

ACL 2023short

In a controllable text generation dataset, there exist unannotated attributes that could provide irrelevant learning signals to models that use it for training and thus degrade their performance. We propose focused prefix tuning (FPT) to mitigate the problem and to enable the control to focus on the…

Cited by 10SourcePDFScholar
2021

Dance Revolution: Long-Term Dance Generation with Music via Curriculum Learning

ICLR 2021poster

Dancing to music is one of human's innate abilities since ancient times. In machine learning research, however, synthesizing dance movements from music is a challenging problem. Recently, researchers synthesize human motion sequences through autoregressive models like recurrent neural network (RNN).…

Cited by 159SourcePDFScholar
2018

Image Recognition Based on Separable Lattice Hmms Using a Deep Neural Network for Output Probability Distributions

ICASSP 2018accepted

This paper proposes an image recognition method based on separable lattice hidden Markov models (SLHMMs) using a deep neural network (DNN) for output probability distributions. The geometric variations of the object to be recognized, e.g., size and location, are essential in image recognition. SLHMM…

Cited by 0SourceScholar
2017

Image recognition based on discriminative models using features generated from separable lattice HMMS

ICASSP 2017accepted

This paper presents an image recognition technique based on discriminative models using features generated from separable lattice hidden Markov models (SL-HMMs). A major problem in image recognition is that the recognition performance is degraded by geometric variations such as that in position and…

Cited by 0SourceScholar