← Search

Satoshi Kobashikawa

4 accepted papers

2024

Talking Face Generation for Impression Conversion Considering Speech Semantics

ICASSP 2024accepted

This study investigates the talking face generation method to convert a speaker’s video to give a target impression, such as “favorable” or “considerate”. Such an impression conversion method needs to consider the input speech semantics because they affect the impression of a speaker’s video along w…

Cited by 0SourceScholar
2023

Next-Speaker Prediction Based on Non-Verbal Information in Multi-Party Video Conversation

ICASSP 2023accepted

We propose a method for next-speaker prediction, a task to predict who speaks in the next turn among multiple current listeners, in multi-party video conversation. Previous studies used non-verbal features, such as head movements and gaze behavior, for next-speaker prediction in face-to-face convers…

Cited by 0SourceScholar
2020

Improving Speaker-Attribute Estimation by Voting Based on Speaker Cluster Information

ICASSP 2020accepted

This paper proposes a general post-processing method for improving speaker-attribute estimation. Estimating speaker-specific attributes such as age and gender is an important task with a wide range of applications. While the recent proposed deep neural network-based end-to-end approach achieves high…

Cited by 0SourceScholar
2018

Soft-Target Training with Ambiguous Emotional Utterances for DNN-Based Speech Emotion Classification

ICASSP 2018accepted

This paper presents a novel emotion classification method for natural speech. One of the problems in the state-of-the-art method based on Deep Neural Network (DNN) is the paucity of the training data compared to model complexity. To solve this problem, this paper utilizes the ambiguous emotional utt…

Cited by 0SourceScholar