← Search

Yukiya Hono

8 accepted papers

2024

Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition

ACL 2024findings

Advances in machine learning have made it possible to perform various text and speech processing tasks, such as automatic speech recognition (ASR), in an end-to-end (E2E) manner. E2E approaches utilizing pre-trained models are gaining attention for conserving training data and resources. However, mo…

2024

PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems

EMNLP 2024finding

Multimodal language models that process both text and speech have a potential for applications in spoken dialogue systems. However, current models face two major challenges in response generation latency: (1) generating a spoken response requires the prior generation of a written response, and (2) s…

2024

PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic Model

ICASSP 2024accepted

This paper presents a neural vocoder based on a denoising diffusion probabilistic model (DDPM) incorporating explicit periodic signals as auxiliary conditioning signals. Recently, DDPM-based neural vocoders have gained prominence as non-autoregressive models that can generate high-quality waveforms.…

Cited by 0SourceScholar
2024

Release of Pre-Trained Models for the Japanese Language

COLING 2024main

AI democratization aims to create a world in which the average person can utilize AI techniques. To achieve this goal, numerous research institutes have attempted to make their results accessible to the public. In particular, large pre-trained models trained on large-scale data have shown unpreceden…

Cited by 18SourcePDFScholar
2023

Embedding a Differentiable Mel-Cepstral Synthesis Filter to a Neural Speech Synthesis System

ICASSP 2023accepted

This paper integrates a classic mel-cepstral synthesis filter into a modern neural speech synthesis system towards end-to-end controllable speech synthesis. Since the mel-cepstral synthesis filter is explicitly embedded in neural waveform models in the proposed system, both voice characteristics and…

Cited by 0SourceScholar
2023

Singing Voice Synthesis Based on a Musical Note Position-Aware Attention Mechanism

ICASSP 2023accepted

This paper proposes a novel sequence-to-sequence (seq2seq) model with a musical note position-aware attention mechanism for singing voice synthesis (SVS). A seq2seq modeling approach that can simultaneously perform acoustic and temporal modeling is attractive. However, due to the difficulty of the t…

Cited by 0SourceScholar
2021

Periodnet: A Non-Autoregressive Waveform Generation Model with a Structure Separating Periodic and Aperiodic Components

ICASSP 2021accepted

We propose PeriodNet, a non-autoregressive (non-AR) waveform generation model with a new model structure for modeling periodic and aperiodic components in speech waveforms. The non-AR waveform generation models can generate speech waveforms parallelly and can be used as a speech vocoder by condition…

Cited by 0SourceScholar
2019

Singing Voice Synthesis Based on Generative Adversarial Networks

ICASSP 2019accepted

This paper proposes a generative adversarial training method for deep neural network (DNN)-based singing voice synthesis. The DNN-based approach has been used in statistical parametric singing voice synthesis and improved the naturalness of the synthesized singing voice [1]. Recently, generative adv…

Cited by 0SourceScholar