← Search

Katsuya Takanashi

2 accepted papers

2025

Do Multimodal Large Language Models Truly See What We Point At? Investigating Indexical, Iconic, and Symbolic Gesture Comprehension

ACL 2025short

Understanding hand gestures is essential for human communication, yet it remains unclear how well multimodal large language models (MLLMs) comprehend them. In this paper, we examine MLLMs’ ability to interpret indexical gestures, which require external referential grounding, in comparison to iconic…

Cited by 0SourcePDFScholar
2018

Audio-Visual Conversation Analysis by Smart Posterboard and Humanoid Robot

ICASSP 2018accepted

This paper addresses audio-visual signal processing for conversation analysis, which involves multi-modal behavior detection and mental-state recognition. We have investigated prediction of turn-taking by the audience in a poster session from their multi-modal behaviors, and found out that the eye-g…

Cited by 0SourceScholar