2025
Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction
ACL 2025finding
Turn-taking is richly multimodal. Predictive turn-taking models (PTTMs) facilitate natural- istic human-robot interaction, yet most rely solely on speech. We introduce MM-VAP, a multimodal PTTM which combines speech with visual cues including facial expression, head pose and gaze. We find that it ou…