← Search

Keito Kudo

4 accepted papers

2026

SLVMEval: Synthetic Meta Evaluation Benchmark for Text-to-Long Video Generation

CVPR 2026

This paper proposes the synthetic long-video meta-evaluation (SLVMEval), a benchmark to perform meta-evaluations of text-to-video (T2V) evaluation systems. The proposed SLVMEval benchmark focuses on assessing these systems on videos of up to 10,486 s (approximately 3 h). The benchmark targets a fund

Cited by 0SourceScholar
2025

Weight-based Analysis of Detokenization in Language Models: Understanding the First Stage of Inference Without Inference

NAACL 2025findings

According to the stages-of-inference hypothesis, early layers of language models map their subword-tokenized input, which does not necessarily correspond to a linguistically meaningful segmentation, to more meaningful representations that form the model’s “inner vocabulary”.Prior analysis of this *d…

2024

First Heuristic Then Rational: Dynamic Use of Heuristics in Language Model Reasoning

EMNLP 2024main

Explicit multi-step reasoning, such as chain-of-thought, is widely adopted in the community to explore the better performance of language models (LMs). We report on the systematic strategy that LMs use in this process.Our controlled experiments reveal that LMs rely more heavily on heuristics, such a…

2023

A Challenging Multimodal Video Summary: Simultaneously Extracting and Generating Keyframe-Caption Pairs from Video

EMNLP 2023long main

This paper proposes a practical multimodal video summarization task setting and a dataset to train and evaluate the task. The target task involves summarizing a given video into a predefined number of keyframe-caption pairs and displaying them in a listable format to grasp the video content quickly.…

Cited by 0SourcecodeScholar