← Search

Koki Maeda

7 accepted papers

2026

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models

ICML 2026poster

While multimodal large language models (MLLMs) have made substantial progress in single-image spatial reasoning, multi-image spatial reasoning, which requires integration of information from multiple viewpoints, remains challenging. Cognitive studies suggest that humans address such tasks through tw…

Cited by 0SourceScholar
2025

Constructing Multimodal Datasets from Scratch for Rapid Development of a Japanese Visual Language Model

NAACL 2025system demonstrations

To develop high-performing Visual Language Models (VLMs), it is essential to prepare multimodal resources, such as image-text pairs, interleaved data, and instruction data. While multimodal resources for English are abundant, there is a significant lack of corresponding resources for non-English lan…

2025

LegalViz: Legal Text Visualization by Text To Diagram Generation

NAACL 2025long

Legal documents including judgments and court orders require highly sophisticated legal knowledge for understanding. To disclose expert knowledge for non-experts, we explore the problem of visualizing legal texts with easy-to-understand diagrams and propose a novel dataset of LegalViz with 23 langua…

2024

COM Kitchens: An Unedited Overhead-view Procedural Videos Dataset a Vision-Language Benchmark

ECCV 2024poster

"Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works often use web videos as training resources, making it challenging to query instructional contents from raw video observa…

2023

DueT: Image-Text Contrastive Transfer Learning with Dual-adapter Tuning

EMNLP 2023long main

This paper presents DueT, a novel transfer learning method for vision and language models built by contrastive learning. In DueT, adapters are inserted into the image and text encoders, which have been initialized using models pre-trained on uni-modal corpora and then frozen. By training only these…

Cited by 0SourceScholar
2023

Query-based Image Captioning from Multi-context 360° Images

EMNLP 2023long findings

A 360-degree image captures the entire scene without the limitations of a camera's field of view, which makes it difficult to describe all the contexts in a single caption. We propose a novel task called Query-based Image Captioning (QuIC) for 360-degree images, where a query (words or short phrases…

Cited by 1SourceScholar