← Search

Seitaro Otsuki

5 accepted papers

2026

LLM-Free Image Captioning Evaluation in Reference-Flexible Settings

AAAI 2026technical

We focus on the automatic evaluation of image captions in both reference-based and reference-free settings. Existing metrics based on large language models (LLMs) favor their own generations; therefore, the neutrality is in question. Most LLM-free metrics do not suffer from such an issue, whereas th

Cited by 0SourcePDFScholar
2025

VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions

EMNLP 2025

In this study, we focus on the automatic evaluation of long and detailed image captions generated by multimodal Large Language Models (MLLMs). Most existing automatic evaluation metrics for image captioning are primarily designed for short captions and are not suitable for evaluating long captions.

Cited by 0SourcePDFScholar
2024

Task Success Prediction for Open-Vocabulary Manipulation Based on Multi-Level Aligned Representations

CoRL 2024poster

In this study, we consider the problem of predicting task success for open-vocabulary manipulation by a manipulator, based on instruction sentences and egocentric images before and after manipulation. Conventional approaches, including multimodal large language models (MLLMs), often fail to appropri…

Cited by 2SourceScholar
2023

Prototypical Contrastive Transfer Learning for Multimodal Language Understanding

IROS 2023poster

Although domestic service robots are expected to assist individuals who require support, they cannot currently interact smoothly with people through natural language. For example, given the instruction “Bring me a bottle from the kitchen,” it is difficult for such robots to specify the bottle in an…

Cited by 2SourceScholar
2022

Shared Transformer Encoder with Mask-Based 3d Model Estimation for Container Mass Estimation

ICASSP 2022accepted

For human-safe robot control in human-to-robot handover, the physical properties of containers and fillings should be accurately estimated. In this paper, we propose a Transformer encoder that shares the same architecture and parameters for filling level and type estimation. We also propose a mask-b…

Cited by 0SourceScholar