← Search

Tomohiro Tanaka

16 accepted papers

2026

Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models

AAAI 2026technical

Contrastive pre-trained vision-language models, such as CLIP, demonstrate strong generalization abilities in zero-shot classification by leveraging embeddings extracted from image and text encoders. This paper aims to robustly fine-tune these vision-language models on in-distribution (ID) data witho

Cited by 0SourcePDFScholar
2025

Multimodal Fine-Grained Apparent Personality Trait Recognition: Joint Modeling of Big Five and Questionnaire Item-level Scores

AAAI 2025technical

This paper presents a novel method for automatically recognizing people's apparent personality traits as perceived by others. In previous studies, apparent personality trait recognition from multimodal human behavior is often modeled to directly estimate personality trait scores, i.e., the ``Big Fiv…

Cited by 0SourcePDFScholar
2024

Talking Face Generation for Impression Conversion Considering Speech Semantics

ICASSP 2024accepted

This study investigates the talking face generation method to convert a speaker’s video to give a target impression, such as “favorable” or “considerate”. Such an impression conversion method needs to consider the input speech semantics because they affect the impression of a speaker’s video along w…

Cited by 0SourceScholar
2023

Exploration of Language Dependency for Japanese Self-Supervised Speech Representation Models

ICASSP 2023accepted

Self-supervised learning (SSL) has been dramatically successful not only in monolingual but also in cross-lingual settings. However, since the two settings have been studied individually in general, there has been little research focusing on how effective a cross-lingual model is in comparison with…

Cited by 5SourceScholar
2023

Improving Scheduled Sampling for Neural Transducer-Based ASR

ICASSP 2023accepted

The recurrent neural network-transducer (RNNT) is a promising approach for automatic speech recognition (ASR) with the introduction of a prediction network that autoregressively considers linguistic aspects. To train the autoregressive part, the ground-truth tokens are used as substitutions for the…

Cited by 0SourceScholar
2023

Leveraging Language Embeddings for Cross-Lingual Self-Supervised Speech Representation Learning

ICASSP 2023accepted

In this paper, we propose novel cross-lingual self-supervised speech representation learning methods that explicitly consider language information. Cross-lingual self-supervised speech representation learning has been studied to make effective use of diverse data in various languages. Previous metho…

Cited by 0SourceScholar
2023

Leveraging Large Text Corpora For End-To-End Speech Summarization

ICASSP 2023accepted

End-to-end speech summarization (E2E SSum) is a technique to directly generate summary sentences from speech. Compared with the cascade approach, which combines automatic speech recognition (ASR) and text summarization models, the E2E approach is more promising because it mitigates ASR errors, incor…

Cited by 0SourceScholar
2022

Hybrid RNN-T/Attention-Based Streaming ASR with Triggered Chunkwise Attention and Dual Internal Language Model Integration

ICASSP 2022accepted

In this paper we propose improvements to our recently proposed hybrid RNN-T/Attention architecture that includes a shared encoder followed by recurrent neural network-transducer (RNN-T) and triggered attention-based decoders (TAD). The use of triggered attention enables the attention-based decoder (…

Cited by 0SourceScholar
2021

Audio-Visual Speech Separation Using Cross-Modal Correspondence Loss

ICASSP 2021accepted

We present an audio-visual speech separation learning method that considers the correspondence between the separated signals and the visual signals to reflect the speech characteristics during training. Audio-visual speech separation is a technique to estimate the individual speech signals from a mi…

Cited by 0SourceScholar
2021

Hierarchical Transformer-Based Large-Context End-To-End ASR with Large-Context Knowledge Distillation

ICASSP 2021accepted

We present a novel large-context end-to-end automatic speech recognition (E2E-ASR) model and its effective training method based on knowledge distillation. Common E2E-ASR models have mainly focused on utterance-level processing in which each utterance is independently transcribed. On the other hand,…

Cited by 0SourceScholar
2021

MAPGN: Masked Pointer-Generator Network for Sequence-to-Sequence Pre-Training

ICASSP 2021accepted

This paper presents a self-supervised learning method for pointer-generator networks to improve spoken-text normalization. Spoken-text normalization that converts spoken-style text into style normalized text is becoming an important technology for improving subsequent processing such as machine tran…

Cited by 0SourceScholar
2021

Simpleflat: A Simple Whole-Network Pre-Training Approach for RNN Transducer-Based End-to-End Speech Recognition

ICASSP 2021accepted

Recurrent neural network-transducer (RNN-T) is promising for building time-synchronous end-to-end automatic speech recognition (ASR) systems, in part because it does not need frame-wise alignment between input features and target labels in the training step. Although training without alignment is be…

Cited by 8SourceScholar
2020

Distilling Attention Weights for CTC-Based ASR Systems

ICASSP 2020accepted

We present a novel training approach for connectionist temporal classification (CTC) -based automatic speech recognition (ASR) systems. CTC models are promising for building both a conventional acoustic model and an end-to-end (E2E) ASR model. However, CTC models make it difficult to capture the cor…

Cited by 0SourceScholar
2020

Spoken Language Acquisition Based on Reinforcement Learning and Word Unit Segmentation

ICASSP 2020accepted

The process of spoken-language acquisition has been one of the topics of greatest interest to linguists for decades. By uti-lizing modern machine learning techniques, we simulated this process on computers, which helps to understand it and develop new possibilities of applying this concept on intell…

Cited by 0SourceScholar
2019

Large Context End-to-end Automatic Speech Recognition via Extension of Hierarchical Recurrent Encoder-decoder Models

ICASSP 2019accepted

This paper describes a novel end-to-end automatic speech recognition (ASR) method that takes into consideration long-range sequential context information beyond utterance boundaries. In spontaneous ASR tasks such as those for discourses and conversations, the input speech often comprises a series of…

Cited by 0SourceScholar