← Search

Taichi Asami

8 accepted papers

2024

What Do Self-Supervised Speech and Speaker Models Learn? New Findings from a Cross Model Layer-Wise Analysis

ICASSP 2024accepted

Self-supervised learning (SSL) has attracted increased attention for learning meaningful speech representations. Speech SSL models, such as WavLM, employ masked prediction training to encode general-purpose representations. In contrast, speaker SSL models, exemplified by DINO-based models, adopt utt…

Cited by 0SourceScholar
2023

An Improved Approximation Algorithm for Wage Determination and Online Task Allocation in Crowd-Sourcing

AAAI 2023technical

Crowd-sourcing has attracted much attention due to its growing importance to society, and numerous studies have been conducted on task allocation and wage determination. Recent works have focused on optimizing task allocation and workers' wages, simultaneously. However, existing methods do not provi…

Cited by 4SourcePDFScholar
2022

Fast Bayesian Estimation of Point Process Intensity as Function of Covariates

NeurIPS 2022accept

In this paper, we tackle the Bayesian estimation of point process intensity as a function of covariates. We propose a novel augmentation of permanental process called augmented permanental process, a doubly-stochastic point process that uses a Gaussian process on covariate space to describe the Baye…

Cited by 5SourcePDFScholar
2018

Neural Confnet Classification: Fully Neural Network Based Spoken Utterance Classification Using Word Confusion Networks

ICASSP 2018accepted

This paper describes neural ConfNet classification, a novel fully neural network based spoken utterance classification method that uses word confusion networks (ConfNets). Our motivation is to establish a spoken utterance classification method that can precisely understand natural language and robus…

Cited by 0SourceScholar
2017

Cross-modal transfer with neural word vectors for image feature learning

ICASSP 2017accepted

Neural word vector (NWV) such as word2vec is a powerful text representation tool that can encode extensive semantic information into compact vectors. This ability poses an interesting question in relation to image processing research - Can we learn better semantic image features from NWVs? We empiri…

Cited by 0SourceScholar
2017

Cumulative moving averaged bottleneck speaker vectors for online speaker adaptation of CNN-based acoustic models

ICASSP 2017accepted

Adapting acoustic models to speakers have shown to greatly improve performance for many tasks. Among the adaptation approaches, exploiting auxiliary features characterizing speakers or environments has received great attention because they allow rapid adaptation, i.e. adaptation with limited amount…

Cited by 0SourceScholar
2017

Domain adaptation of DNN acoustic models using knowledge distillation

ICASSP 2017accepted

Constructing deep neural network (DNN) acoustic models from limited training data is an important issue for the development of automatic speech recognition (ASR) applications that will be used in various application-specific acoustic environments. To this end, domain adaptation techniques that train…

Cited by 0SourceScholar
2017

Parallel phonetically aware DNNs and LSTM-RNNS for frame-by-frame discriminative modeling of spoken language identification

ICASSP 2017accepted

Parallel phonetically aware deep neural networks (PPA-DNNs) and long short-term memory recurrent neural networks (PPA-LSTM-RNNs) to enhance frame-by-frame discriminative modeling of spoken language identification are proposed. This idea is inspired by traditional systems based on parallel phoneme re…

Cited by 0SourceScholar