← Search

Koichiro Yoshino

10 accepted papers

2025

Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures

ACL 2025long

Multimodal reference resolution, including phrase grounding, aims to understand the semantic relations between mentions and real-world objects. Phrase grounding between images and their captions is a well-established task. In contrast, for real-world applications, it is essential to integrate textua…

2025

J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception

IROS 2025

We introduce J-ORA, a novel multimodal dataset that bridges the gap in robot perception by providing detailed object attribute annotations within Japanese human-robot dialogue scenarios. J-ORA is designed to support three critical perception tasks, object identification, reference resolution, and ne

Cited by 0SourcecodeScholar
2025

Proactive User Information Acquisition via Chats on User-Favored Topics

EMNLP 2025

Chat-oriented dialogue systems that deliver tangible benefits, such as sharing news or frailty prevention for seniors, require proactive acquisition of specific user information via chats on user-favored topics. This study proposes the Proactive Information Acquisition (PIA) task to support the deve

2025

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression

ICASSP 2025accepted

To improve user engagement during conversations with dialogue systems, we must improve individual dialogue responses and dialogue impressions such as consistency, personality, and empathy throughout the entire dialogue. While such dialogue systems have been developing rapidly with the help of large…

Cited by 0SourceScholar
2024

A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions

COLING 2024main

Situated conversations, which refer to visual information as visual question answering (VQA), often contain ambiguities caused by reliance on directive information. This problem is exacerbated because some languages, such as Japanese, often omit subjective or objective terms. Such ambiguities in que…

2024

J-CRe3: A Japanese Conversation Dataset for Real-world Reference Resolution

COLING 2024main

Understanding expressions that refer to the physical world is crucial for such human-assisting systems in the real world, as robots that must perform actions that are expected by users. In real-world reference resolution, a system must ground the verbal information that appears in user interactions…

2023

Analysis of Style-Shifting on Social Media: Using Neural Language Model Conditioned by Social Meanings

EMNLP 2023long findings

In this paper, we propose a novel framework for evaluating style-shifting in social media conversations. Our proposed framework captures changes in an individual's conversational style based on surprisals predicted by a personalized neural language model for individuals. Our personalized language mo…

Cited by 0SourceScholar
2023

Operative Action Captioning for Estimating System Actions

ICRA 2023poster

Human-assistive systems, such as robots, need to correctly understand the surrounding situation based on obser-vations and output the required support actions for humans. Language is one of the important channels to communicate with humans, and robots are required to have the ability to express thei…

Cited by 0SourceScholar
2020

Improving Spoken Language Understanding by Wisdom of Crowds

COLING 2020main

Spoken language understanding (SLU), which converts user requests in natural language to machine-interpretable expressions, is becoming an essential task. The lack of training data is an important problem, especially for new system tasks, because existing SLU systems are based on statistical approac…