← Search

Yuiga Wada

6 accepted papers

2026

LLM-Free Image Captioning Evaluation in Reference-Flexible Settings

AAAI 2026technical

We focus on the automatic evaluation of image captions in both reference-based and reference-free settings. Existing metrics based on large language models (LLMs) favor their own generations; therefore, the neutrality is in question. Most LLM-free metrics do not suffer from such an issue, whereas th

Cited by 0SourcePDFScholar
2026

ZINA: Multimodal Fine-grained Hallucination Detection and Editing

CVPR 2026

Multimodal Large Language Models (MLLMs) often generate hallucinations, where the output deviates from the visual content. Given that these hallucinations can take diverse forms, detecting hallucinations at a fine-grained level is essential for comprehensive evaluation and analysis. To this end, we

Cited by 0SourcecodeScholar
2025

VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions

EMNLP 2025

In this study, we focus on the automatic evaluation of long and detailed image captions generated by multimodal Large Language Models (MLLMs). Most existing automatic evaluation metrics for image captioning are primarily designed for short captions and are not suitable for evaluating long captions.

Cited by 0SourcePDFScholar
2024

Polos: Multimodal Metric Learning from Human Feedback for Image Captioning

CVPR 2024highlight

Establishing an automatic evaluation metric that closely aligns with human judgments is essential for effectively developing image captioning models. Recent data-driven metrics have demonstrated a stronger correlation with human judgments than classic metrics such as CIDEr; however they lack suffici…

2023

Multimodal Diffusion Segmentation Model for Object Segmentation from Manipulation Instructions

IROS 2023poster

In this study, we aim to develop a model that comprehends a natural language instruction (e.g., “Go to the living room and get the nearest pillow to the radio art on the wall”) and generates a segmentation mask for the target everyday object. The task is challenging because it requires (1) the under…

Cited by 6SourceScholar
2022

Shared Transformer Encoder with Mask-Based 3d Model Estimation for Container Mass Estimation

ICASSP 2022accepted

For human-safe robot control in human-to-robot handover, the physical properties of containers and fillings should be accurately estimated. In this paper, we propose a Transformer encoder that shares the same architecture and parameters for filling level and type estimation. We also propose a mask-b…

Cited by 0SourceScholar