← Search

Koji Inoue

7 accepted papers

2025

A Noise-Robust Turn-Taking System for Real-World Dialogue Robots: A Field Experiment

IROS 2025

Turn-taking is a crucial aspect of human-robot interaction, directly influencing conversational fluidity and user engagement. While previous research has explored turn-taking models in controlled environments, their robustness in real-world settings remains underexplored. In this study, we propose a

Cited by 3SourceScholar
2025

Do Multimodal Large Language Models Truly See What We Point At? Investigating Indexical, Iconic, and Symbolic Gesture Comprehension

ACL 2025short

Understanding hand gestures is essential for human communication, yet it remains unclear how well multimodal large language models (MLLMs) comprehend them. In this paper, we examine MLLMs’ ability to interpret indexical gestures, which require external referential grounding, in comparison to iconic…

Cited by 0SourcePDFScholar
2025

Human-Like Embodied AI Interviewer: Employing Android ERICA in Real International Conference

COLING 2025system demonstrations

This paper introduces the human-like embodied AI interviewer which integrates android robots equipped with advanced conversational capabilities, including attentive listening, conversational repairs, and user fluency adaptation. Moreover, it can analyze and present results post-interview. We conduct…

2025

Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection

NAACL 2025long

In human conversations, short backchannel utterances such as “yeah” and “oh” play a crucial role in facilitating smooth and engaging dialogue.These backchannels signal attentiveness and understanding without interrupting the speaker, making their accurate prediction essential for creating more natur…

Cited by 2SourcePDFScholar
2024

Multilingual Turn-taking Prediction Using Voice Activity Projection

COLING 2024main

This paper investigates the application of voice activity projection (VAP), a predictive turn-taking model for spoken dialogue, on multilingual data, encompassing English, Mandarin, and Japanese. The VAP model continuously predicts the upcoming voice activities of participants in dyadic dialogue, le…

Cited by 9SourcePDFScholar
2018

An End-to-End Approach to Joint Social Signal Detection and Automatic Speech Recognition

ICASSP 2018accepted

Social signals such as laughter and fillers are often observed in natural conversation, and they play various roles in human-to-human communication. Detecting these events is useful for transcription systems to generate rich transcription and for dialogue systems to behave as we do such as synchroni…

Cited by 0SourceScholar
2018

Audio-Visual Conversation Analysis by Smart Posterboard and Humanoid Robot

ICASSP 2018accepted

This paper addresses audio-visual signal processing for conversation analysis, which involves multi-modal behavior detection and mental-state recognition. We have investigated prediction of turn-taking by the audience in a poster session from their multi-modal behaviors, and found out that the eye-g…

Cited by 0SourceScholar