← Search

Kazuki Hayashi

5 accepted papers

2025

BQA: Body Language Question Answering Dataset for Video Large Language Models

ACL 2025short

A large part of human communication relies on nonverbal cues such as facial expressions, eye contact, and body language. Unlike language or sign language, such nonverbal communication lacks formal rules, requiring complex reasoning based on commonsense understanding.Enabling current Video Large Lang…

Cited by 0SourcePDFScholar
2025

IRR: Image Review Ranking Framework for Evaluating Vision-Language Models

COLING 2025main

Large-scale Vision-Language Models (LVLMs) process both images and text, excelling in multimodal tasks such as image captioning and description generation. However, while these models excel at generating factual content, their ability to generate and evaluate texts reflecting perspectives on the sam…

Cited by 1SourcePDFScholar
2025

LoCt-Instruct: An Automatic Pipeline for Constructing Datasets of Logical Continuous Instructions

EMNLP 2025

Continuous instruction following closely mirrors real-world tasks by requiring models to solve sequences of interdependent steps, yet existing multi-step instruction datasets suffer from three key limitations: (1) lack of logical coherence across turns, (2) narrow topical breadth and depth, and (3)

2025

Towards Cross-Lingual Explanation of Artwork in Large-scale Vision Language Models

NAACL 2025findings

As the performance of Large-scale Vision Language Models (LVLMs) improves, they are increasingly capable of responding in multiple languages, and there is an expectation that the demand for explanations generated by LVLMs will grow. However, pre-training of Vision Encoder and the integrated training…

Cited by 5SourcePDFScholar
2024

Towards Artwork Explanation in Large-scale Vision Language Models

ACL 2024short

Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating advanced capabilities in text generation and comprehension. However, it has not been clarified to what extent LVLMs understand the knowledge necessary for explaining images, the complex relationships b…