← Search

Takumi Hirose

4 accepted papers

2026

DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning

AAAI 2026technical

Large vision-language models (LVLMs) have shown impressive performance across a broad range of multimodal tasks. However, robust image caption evaluation using LVLMs remains challenging, particularly under domain-shift scenarios. To address this issue, we introduce the Distribution-Aware Score Decod

Cited by 0SourcePDFScholar
2025

HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

ICLR 2025poster

Recently, Text-to-speech (TTS) models based on large language models (LLMs) that translate natural language text into sequences of discrete audio tokens have gained great research attention, with advances in neural audio codec (NAC) mod- els using residual vector quantization (RVQ). However, long-fo…

2025

Multi-Point Positional Insertion Tuning for Small Object Detection

ICASSP 2025accepted

Small object detection aims to localize and classify small objects within images. With recent advances in large-scale vision-language pretraining, finetuning pretrained object detection models has emerged as a promising approach. However, finetuning large models is computationally and memory expensi…

Cited by 0SourceScholar
2025

Referring Expression Comprehension for Small Objects

ICCV 2025poster

Referring expression comprehension (REC) aims to localize the target object described by a natural language expression.Recent advances in vision-language learning have led to significant performance improvements in REC tasks.However, localizing extremely small objects remains a considerable challeng…