← Search

Satoshi Nakamura

27 accepted papers

2026

SASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech Translation

AAAI 2026technical

This work proposes a grammar-based chunking strategy that segments input streams into semantically complete units by parsing dependency relations (e.g., noun phrase boundaries, verb-object structures) and punctuation features. The method ensures chunk coherence and minimizes semantic fragmentation.

Cited by 0SourcePDFScholar
2024

LLMs Are Zero-Shot Context-Aware Simultaneous Translators

EMNLP 2024main

The advent of transformers has fueled progress in machine translation. More recently large language models (LLMs) have come to the spotlight thanks to their generality and strong performance in a wide range of language tasks, including translation. Here we show that open-source LLMs perform on par w…

2024

LLaST: Improved End-to-end Speech Translation System Leveraged by Large Language Models

ACL 2024findings

We introduces ***LLaST***, a framework for building high-performance Large Language model based Speech-to-text Translation systems. We address the limitations of end-to-end speech translation (E2E ST) models by exploring model architecture design and optimization techniques tailored for LLMs. Our ap…

2024

NAIST-SIC-Aligned: An Aligned English-Japanese Simultaneous Interpretation Corpus

COLING 2024main

It remains a question that how simultaneous interpretation (SI) data affects simultaneous machine translation (SiMT). Research has been limited due to the lack of a large-scale training corpus. In this work, we aim to fill in the gap by introducing NAIST-SIC-Aligned, which is an automatically-aligne…

2024

Subspace Representations for Soft Set Operations and Sentence Similarities

NAACL 2024long

In the field of natural language processing (NLP), continuous vector representations are crucial for capturing the semantic meanings of individual words. Yet, when it comes to the representations of sets of words, the conventional vector-based approaches often struggle with expressiveness and lack t…

2024

TransLLaMa: LLM-based Simultaneous Translation System

EMNLP 2024finding

Decoder-only large language models (LLMs) have recently demonstrated impressive capabilities in text generation and reasoning. Nonetheless, they have limited applications in simultaneous machine translation (SiMT), currently dominated by encoder-decoder transformers. This study demonstrates that, af…

2023

Self-Adaptive Incremental Machine Speech Chain for Lombard TTS with High-Granularity ASR Feedback in Dynamic Noise Condition

ICASSP 2023accepted

A common approach for text-to-speech (TTS) in noisy conditions is offline fine-tuning, which is generally utilized on static noises and predefined conditions. We recently proposed a self-adaptive TTS in machine speech chain inference that enables TTS to control its voices in statically and dynamical…

Cited by 0SourceScholar
2022

USB: A Unified Semi-supervised Learning Benchmark for Classification

NeurIPS 2022accept

Semi-supervised learning (SSL) improves model generalization by leveraging massive unlabeled data to augment limited labeled samples. However, currently, popular SSL evaluation protocols are often constrained to computer vision (CV) tasks. In addition, previous work typically trains deep neural netw…

2020

DEJA-VU: Double Feature Presentation and Iterated Loss in Deep Transformer Networks

ICASSP 2020accepted

Deep acoustic models typically receive features in the first layer of the network, and process increasingly abstract representations in the subsequent layers. Here, we propose to feed the input features at multiple depths in the acoustic model. As our motivation is to allow acoustic models to re-exa…

Cited by 0SourceScholar
2020

Improving Spoken Language Understanding by Wisdom of Crowds

COLING 2020main

Spoken language understanding (SLU), which converts user requests in natural language to machine-interpretable expressions, is becoming an essential task. The lack of training data is an important problem, especially for new system tasks, because existing SLU systems are based on statistical approac…

2020

Incorporating Noisy Length Constraints into Transformer with Length-aware Positional Encodings

COLING 2020main

Neural Machine Translation often suffers from an under-translation problem due to its limited modeling of output sequence lengths. In this work, we propose a novel approach to training a Transformer model using length constraints based on length-aware positional encoding (PE). Since length constrain…

Cited by 11SourcePDFScholar
2020

Using Panoramic Videos for Multi-Person Localization and Tracking In A 3D Panoramic Coordinate

ICASSP 2020accepted

3D panoramic multi-person localization and tracking are prominent in many applications, however, conventional methods using LiDAR equipment could be economically expensive and also computationally inefficient due to the processing of point cloud data. In this work, we propose an effective and effici…

Cited by 0SourceScholar
2019

Cross-lingual Speech-based Tobi Label Generation Using Bidirectional Lstm

ICASSP 2019accepted

In this paper we investigate the automatic generation of ToBI-style prosody labels. The work is motivated by the idea of using prosodic information to facilitate the automatic lexicon discovery for unseen and under-resourced languages for which sufficient training data is not available. Specifically…

Cited by 0SourceScholar
2019

End-to-end Feedback Loss in Speech Chain Framework via Straight-through Estimator

ICASSP 2019accepted

The speech chain mechanism integrates automatic speech recognition (ASR) and text-to-speech synthesis (TTS) modules into a single cycle during training. In our previous work, we applied a speech chain mechanism as a semi-supervised learning. It provides the ability for ASR and TTS to assist each oth…

Cited by 0SourceScholar
2019

Speech Artifact Removal from Eeg Recordings of Spoken Word Production with Tensor Decomposition

ICASSP 2019accepted

Research about brain activities involving spoken word production is considerably underdeveloped because of the undiscovered characteristics of speech artifacts, which contaminate electroencephalogram (EEG) signals and prevent the inspection of the underlying cognitive processes. To fuel further EEG…

Cited by 0SourceScholar
2018

Graph Regularized Tensor Factorization for Single-Trial EEG Analysis

ICASSP 2018accepted

This study proposes a tensor factorization algorithm for electroencephalographies (EEGs) that incorporates the geometric structure of the electrode location. The purpose is removing noise caused by EEG activities which are irrelevant to stimuli presented to a subject from single-trial event-related…

Cited by 0SourceScholar
2016

An estimation method of voice timbre evaluation values using feature extraction with Gaussian mixture model based on reference singer

ICASSP 2016accepted

This paper presents an estimation method of voice timbre evaluation values for arbitrary singer's singing voices generated with a singing voice synthesis system towards the development of a singing voice retrieval system. The voice timbre evaluation values are numerical values corresponding to voice…

Cited by 0SourceScholar
2016

Implementation of F0 transformation for statistical singing voice conversion based on direct waveform modification

ICASSP 2016accepted

This paper presents a technique for transforming F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> in a framework of statistical singing voice conversion with direct waveform modification based on spectrum differential (DIFFSVC). The DIFFSVC met…

Cited by 0SourceScholar
2016

Noise suppression method for body-conducted soft speech enhancement based on external noise monitoring

ICASSP 2016accepted

This paper presents a novel approach to suppressing adverse effects of external noise on body-conducted soft speech for silent speech communication in noisy environments. Nonaudible murmur (NAM) microphone as one of the body-conductive microphones is capable of detecting very soft speech. However, b…

Cited by 0SourceScholar
2016

Statistical F0 prediction for electrolaryngeal speech enhancement considering generative process of F0 contours within product of experts framework

ICASSP 2016accepted

We have previously proposed a statistical fundamental frequency (F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> ) prediction method that makes it possible to predict the underlying F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:x…

Cited by 0SourceScholar
2015

Combination of two-dimensional cochleogram and spectrogram features for deep learning-based ASR

ICASSP 2015accepted

This paper explores the use of auditory features based on cochleograms; two dimensional speech features derived from gammatone filters within the convolutional neural network (CNN) framework. Furthermore, we also propose various possibilities to combine cochleogram features with log-mel filter banks…

Cited by 22SourceScholar
2015

EEG signal enhancement using multi-channel wiener filter with a spatial correlation prior

ICASSP 2015accepted

Event-related potentials (ERPs) of electroencephalogram (EEG) are often used as features for brain machine interfaces or for analysis of brain activities. However, as EEG signals easily suffer from various artifacts, ERPs are often collapsed and hard to observe. There are several attempts at using m…

Cited by 0SourceScholar
2015

Modulation spectrum-constrained trajectory training algorithm for GMM-based Voice Conversion

ICASSP 2015accepted

This paper presents a novel training algorithm for Gaussian Mixture Model (GMM)-based Voice Conversion (VC). One of the advantages of GMM-based VC is computationally efficient conversion processing enabling to achieve real-time VC applications. On the other hand, the quality of the converted speech…

Cited by 0SourceScholar
2015

Parameter generation algorithm considering Modulation Spectrum for HMM-based speech synthesis

ICASSP 2015accepted

This paper proposes a novel parameter generation algorithm for high-quality speech generation in Hidden Markov Model (HMM)-based speech synthesis. One of the biggest issues causing significant quality degradation is the over-smoothing effect often observed in generated parameter trajectories. Global…

Cited by 15SourceScholar
2015

Statistical modeling of binaural signal and its application to binaural source separation

ICASSP 2015accepted

This paper addresses a new statistical model of binaural signals and its application to efficient binaural source separation. Binaural source separation is always required to retain a spatial cue of the separated sound, such as a head-related transfer function (HRTF). However, the direct use of an H…

Cited by 0SourceScholar
2015

WFST-based structural classification integrating dnn acoustic features and RNN language features for speech recognition

ICASSP 2015accepted

This paper proposes a method to train Weighted Finite State Transducer (WFST) based structural classifiers using deep neural network (DNN) acoustic features and recurrent neural network (RNN) language features for speech recognition. Structural classification is an effective approach to achieve high…

Cited by 0SourceScholar