← Search

Nobukatsu Hojo

8 accepted papers

2025

Multimodal Fine-Grained Apparent Personality Trait Recognition: Joint Modeling of Big Five and Questionnaire Item-level Scores

AAAI 2025technical

This paper presents a novel method for automatically recognizing people's apparent personality traits as perceived by others. In previous studies, apparent personality trait recognition from multimodal human behavior is often modeled to directly estimate personality trait scores, i.e., the ``Big Fiv…

Cited by 0SourcePDFScholar
2025

ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind

AAAI 2025technical

Existing Theory of Mind (ToM) benchmarks diverge from real-world scenarios in three aspects: 1) they assess a limited range of mental states such as beliefs, 2) false beliefs are not comprehensively explored, and 3) the diverse personality traits of characters are overlooked. To address these challe…

2024

Talking Face Generation for Impression Conversion Considering Speech Semantics

ICASSP 2024accepted

This study investigates the talking face generation method to convert a speaker’s video to give a target impression, such as “favorable” or “considerate”. Such an impression conversion method needs to consider the input speech semantics because they affect the impression of a speaker’s video along w…

Cited by 0SourceScholar
2023

Next-Speaker Prediction Based on Non-Verbal Information in Multi-Party Video Conversation

ICASSP 2023accepted

We propose a method for next-speaker prediction, a task to predict who speaks in the next turn among multiple current listeners, in multi-party video conversation. Previous studies used non-verbal features, such as head movements and gaze behavior, for next-speaker prediction in face-to-face convers…

Cited by 0SourceScholar
2021

Maskcyclegan-VC: Learning Non-Parallel Voice Conversion with Filling in Frames

ICASSP 2021accepted

Non-parallel voice conversion (VC) is a technique for training voice converters without a parallel corpus. Cycle-consistent adversarial network-based VCs (CycleGAN-VC and CycleGAN-VC2) are widely accepted as benchmark methods. However, owing to their insufficient ability to grasp time-frequency stru…

Cited by 0SourceScholar
2019

ATTS2S-VC: Sequence-to-sequence Voice Conversion with Attention and Context Preservation Mechanisms

ICASSP 2019accepted

This paper describes a method based on a sequence-to-sequence learning (Seq2Seq) with attention and context preservation mechanism for voice conversion (VC) tasks. Seq2Seq has been outstanding at numerous tasks involving sequence modeling such as speech synthesis and recognition, machine translation…

Cited by 115SourceScholar
2019

Cyclegan-VC2: Improved Cyclegan-based Non-parallel Voice Conversion

ICASSP 2019accepted

Non-parallel voice conversion (VC) is a technique for learning the mapping from source to target speech without relying on parallel data. This is an important task, but it has been challenging due to the disadvantages of the training conditions. Recently, CycleGAN-VC has provided a breakthrough and…

Cited by 0SourceScholar
2017

Generative adversarial network-based postfilter for statistical parametric speech synthesis

ICASSP 2017accepted

We propose a postfilter based on a generative adversarial network (GAN) to compensate for the differences between natural speech and speech synthesized by statistical parametric speech synthesis. In particular, we focus on the differences caused by over-smoothing, which makes the sounds muffled. Ove…

Cited by 0SourceScholar