← Search

Emiru Tsunoo

17 accepted papers

2025

Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features

ICASSP 2025accepted

Real-time speech enhancement (SE) is essential to online speech communication. Causal SE models use only the previous context while predicting future information, such as phoneme continuation, may help performing causal SE. The phonetic information is often represented by quantizing latent features…

Cited by 0SourceScholar
2025

ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems

NAACL 2025system demonstrations

Advancements in audio foundation models (FMs) have fueled interest in end-to-end (E2E) spoken dialogue systems, but different web interfaces for each system makes it challenging to compare and contrast them effectively. Motivated by this, we introduce an open-source, user-friendly toolkit designed t…

2025

Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens

ICASSP 2025accepted

In many real-world scenarios, such as meetings, multiple speakers are present with an unknown number of participants, and their utterances often overlap. We address these multi-speaker challenges by a novel attention-based encoder-decoder method augmented with special speaker class tokens obtained b…

Cited by 0SourceScholar
2024

Phoneme-Aware Encoding for Prefix-Tree-Based Contextual ASR

ICASSP 2024accepted

In speech recognition applications, it is important to recognize context-specific rare words, such as proper nouns. Tree-constrained Pointer Generator (TCPGen) has shown promise for this purpose, which efficiently biases such words with a prefix tree. While the original TCPGen relies on grapheme-bas…

Cited by 0SourceScholar
2024

UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions

NAACL 2024long

Recent studies leverage large language models with multi-tasking capabilities, using natural language prompts to guide the model’s behavior and surpassing performance of task-specific models. Motivated by this, we ask: can we build a single model that jointly performs various spoken language underst…

2023

A Study on the Integration of Pipeline and E2E SLU Systems for Spoken Semantic Parsing Toward Stop Quality Challenge

ICASSP 2023accepted

Recently there have been efforts to introduce new benchmark tasks for spoken language understanding (SLU), like semantic parsing. In this paper, we describe our proposed spoken semantic parsing system for the quality track (Track 1) in Spoken Language Understanding Grand Challenge which is part of I…

Cited by 0SourceScholar
2023

E-Branchformer-Based E2E SLU Toward Stop on-Device Challenge

ICASSP 2023accepted

In this paper, we report our team’s study on track 2 of the Spoken Language Understanding Grand Challenge, which is a component of the ICASSP Signal Processing Grand Challenge 2023. The task is intended for on-device processing and involves estimating semantic parse labels from speech using a model…

Cited by 0SourceScholar
2023

Joint Modelling of Spoken Language Understanding Tasks with Integrated Dialog History

ICASSP 2023accepted

Most human interactions occur in the form of spoken conversations where the semantic meaning of a given utterance depends on the context. Each utterance in spoken conversation can be represented by many semantic and speaker attributes, and there has been an interest in building Spoken Language Under…

Cited by 0SourceScholar
2023

Streaming Joint Speech Recognition and Disfluency Detection

ICASSP 2023accepted

Disfluency detection has mainly been solved in a pipeline approach, as post-processing of speech recognition. In this study, we propose Transformer-based encoder-decoder models that jointly solve speech recognition and disfluency detection, which work in a streaming manner. Compared to pipeline appr…

Cited by 0SourceScholar
2023

The Pipeline System of ASR and NLU with MLM-based data Augmentation Toward Stop Low-Resource Challenge

ICASSP 2023accepted

This paper describes our system for the low-resource domain adaptation track (Track 3) in Spoken Language Understanding Grand Challenge, which is a part of ICASSP Signal Processing Grand Challenge 2023. In the track, we adopt a pipeline approach of ASR and NLU. For ASR, we fine-tune Whisper for each…

Cited by 0SourceScholar
2022

Joint Speech Recognition and Audio Captioning

ICASSP 2022accepted

Speech samples recorded in both indoor and outdoor environments are often contaminated with secondary audio sources. Most end-to-end monaural speech recognition systems either remove these background sounds using speech enhancement or train noise-robust models. For better model interpretability and…

Cited by 0SourceScholar
2022

Multi-ACCDOA: Localizing And Detecting Overlapping Sounds From The Same Class With Auxiliary Duplicating Permutation Invariant Training

ICASSP 2022accepted

Sound event localization and detection (SELD) involves identifying the direction-of-arrival (DOA) and the event class. The SELD methods with a class-wise output format make the model predict activities of all sound event classes and corresponding locations. The class-wise methods can output activity…

Cited by 111SourceScholar
2022

Polyphone Disambiguation and Accent Prediction Using Pre-Trained Language Models in Japanese TTS Front-End

ICASSP 2022accepted

Although end-to-end text-to-speech (TTS) models can generate natural speech, challenges still remain when it comes to estimating sentence-level phonetic and prosodic information from raw text in Japanese TTS systems. In this paper, we propose a method for polyphone disambiguation (PD) and accent pre…

Cited by 0SourceScholar
2022

Run-and-Back Stitch Search: Novel Block Synchronous Decoding For Streaming Encoder-Decoder ASR

ICASSP 2022accepted

A streaming style inference of encoder–decoder automatic speech recognition (ASR) systems is important for reducing latency, which is essential for interactive use cases. To this end, we propose a novel blockwise synchronous decoding algorithm with a hybrid approach that combines endpoint prediction…

Cited by 0SourceScholar
2022

Spatial Data Augmentation with Simulated Room Impulse Responses for Sound Event Localization and Detection

ICASSP 2022accepted

Recording and annotating real sound events for a sound event localization and detection (SELD) task is time consuming, and data augmentation techniques are often favored when the amount of data is limited. However, how to augment the spatial information in a dataset, including unlabeled directional…

Cited by 0SourceScholar
2021

Gaussian Kernelized Self-Attention for Long Sequence Data and its Application to CTC-Based Speech Recognition

ICASSP 2021accepted

ISelf-attention (SA) based models have recently achieved significant performance improvements in hybrid and end-to-end automatic speech recognition (ASR) systems owing to their flexible context modeling capability. However, it is also known that the accuracy degrades when applying SA to long sequenc…

Cited by 0SourceScholar
2021

Making Punctuation Restoration Robust and Fast with Multi-Task Learning and Knowledge Distillation

ICASSP 2021accepted

In punctuation restoration, we try to recover the missing punctuation from automatic speech recognition output to improve understandability. Currently, large pre-trained transformers such as BERT set the benchmark on this task but there are two main drawbacks to these models. First, the pre-training…

Cited by 0SourceScholar