← Search

Yuto Nishimura

2 accepted papers

2025

HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

ICLR 2025poster

Recently, Text-to-speech (TTS) models based on large language models (LLMs) that translate natural language text into sequences of discrete audio tokens have gained great research attention, with advances in neural audio codec (NAC) mod- els using residual vector quantization (RVQ). However, long-fo…

2024

Minimax optimality of convolutional neural networks for infinite dimensional input-output problems and separation from kernel methods

ICLR 2024poster

Recent deep learning applications, exemplified by text-to-image tasks, often involve high-dimensional inputs and outputs. While several studies have investigated the function estimation capabilities of deep learning, research on dilated convolutional neural networks (CNNs) has mainly focused on case…

Cited by 1SourcePDFScholar