← Search

Takehiro Moriya

6 accepted papers

2026

Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec

ICASSP 2026poster

Neural audio codec (NAC) is essential for reconstructing high-quality speech signals and generating discrete representations for downstream speech language models. However, ensuring accurate semantic modeling while maintaining high-fidelity reconstruction under ultra-low bitrate constraints remains…

Cited by 0SourcePDFScholar
2025

Stereo Downmix in 3GPP IVAS for EVS Compatibility

ICASSP 2025accepted

The 3GPP IVAS codec specifies an EVS-compatible stereo downmix as one of the key functionalities. This paper describes how this novel active downmix scheme has been devised to achieve high and stable quality from stereo input to EVS encoder/decoder with no additional algorithmic delay. An example of…

Cited by 0SourceScholar
2019

Detecting Attention Shift from Neural Response Based on Beat-frequency-modulated Musical Excerpts

ICASSP 2019accepted

This paper presents a new approach for detecting attention from auditory steady-state responses (ASSR) by using musical excerpts. The feature extraction process for electroencephalogram (EEG) signal is combined with a support vector machine as a binary discriminator. A novel modulation that emphasiz…

Cited by 0SourceScholar
2018

Spectral-Envelope-Based Least Significant Bit Management for Low-Delay Bit-Error-Robust Speech Coding

ICASSP 2018accepted

We have devised a method for bit assignment of quantized frequency spectra aiming at its use in low-delay bit-error-robust speech compression. The proposed method, least significant bit management (LSBM), controls the least significant bits of the spectra based on their envelope to make them represe…

Cited by 0SourceScholar
2017

Shape parameter estimation for generalized-Gaussian-distributed frequency spectra of audio signals

ICASSP 2017accepted

We have devised a method for estimating, from a single frame of audio frequency spectra, a shape parameter of multivariate generalized Gaussian distribution which has variance represented by an all-pole model and no covariance. Based on powered all-pole spectrum estimation (PAPSE), which is an exten…

Cited by 2SourceScholar
2015

Low delay LPC and MDCT-based audio coding in the EVS codec

ICASSP 2015accepted

Speech coders operating in time domain can be extended with a frequency domain mode to improve encoding of music, even though this is challenging at low delay. In such a scenario, the short analysis window limits the benefit of the transform coder, while a delayless switch between the two coders con…

Cited by 0SourceScholar