← Search

Yusuke Sakai

27 accepted papers

2025

AdTEC: A Unified Benchmark for Evaluating Text Quality in Search Engine Advertising

NAACL 2025long

As the fluency of ad texts automatically generated by natural language generation technologies continues to improve, there is an increasing demand to assess the quality of these creatives in real-world setting.We propose **AdTEC**, the first public benchmark to evaluate ad texts from multiple perspe…

2025

BQA: Body Language Question Answering Dataset for Video Large Language Models

ACL 2025short

A large part of human communication relies on nonverbal cues such as facial expressions, eye contact, and body language. Unlike language or sign language, such nonverbal communication lacks formal rules, requiring complex reasoning based on commonsense understanding.Enabling current Video Large Lang…

Cited by 0SourcePDFScholar
2025

BannerBench: Benchmarking Vision Language Models for Multi-Ad Selection with Human Preferences

EMNLP 2025

Web banner advertisements, which are placed on websites to guide users to a targeted landing page (LP), are still often selected manually because human preferences are important in selecting which ads to deliver. To automate this process, we propose a new benchmark, BannerBench, to evaluate the huma

Cited by 0SourcePDFScholar
2025

Dictionaries to the Rescue: Cross-Lingual Vocabulary Transfer for Low-Resource Languages Using Bilingual Dictionaries

ACL 2025finding

Cross-lingual vocabulary transfer plays a promising role in adapting pre-trained language models to new languages, including low-resource languages.Existing approaches that utilize monolingual or parallel corpora face challenges when applied to languages with limited resources.In this work, we propo…

2025

Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding

ACL 2025long

Inference methods play an important role in eliciting the performance of large language models (LLMs). Currently, LLMs use inference methods utilizing generated multiple samples, which can be derived from Minimum Bayes Risk (MBR) Decoding. Previous studies have conducted empirical analyses to clarif…

2025

Funabot-Sleeve: A Wearable Device Employing McKibben Artificial Muscles for Haptic Sensation in the Forearm

RA-L 2025

Haptic feedback systems play a critical role in enriching the user experience in human-robot interaction. However, existing devices designed for evoking haptic sensations often face limitations owing to their low degree of freedom of deformation. In this study, we introduce the Funabot-Sleeve, a hap

Cited by 31SourceScholar
2025

How to Make the Most of LLMs’ Grammatical Knowledge for Acceptability Judgments

NAACL 2025long

The grammatical knowledge of language models (LMs) is often measured using a benchmark of linguistic minimal pairs, where LMs are presented with a pair of acceptable and unacceptable sentences and required to judge which is more acceptable. Conventional approaches compare sentence probabilities dire…

2025

IMPARA-GED: Grammatical Error Detection is Boosting Reference-free Grammatical Error Quality Estimator

ACL 2025finding

We propose IMPARA-GED, a novel reference-free automatic grammatical error correction (GEC) evaluation method with grammatical error detection (GED) capabilities. We focus on the quality estimator of IMPARA, an existing automatic GEC evaluation method, and construct that of IMPARA-GED using a pre-tra…

2025

IRR: Image Review Ranking Framework for Evaluating Vision-Language Models

COLING 2025main

Large-scale Vision-Language Models (LVLMs) process both images and text, excelling in multimodal tasks such as image captioning and description generation. However, while these models excel at generating factual content, their ability to generate and evaluate texts reflecting perspectives on the sam…

Cited by 1SourcePDFScholar
2025

Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark

EMNLP 2025

Multilingual machine translation (MT) benchmarks play a central role in evaluating the capabilities of modern MT systems. Among them, the FLORES+ benchmark is widely used, offering English-to-many translation data for over 200 languages, curated with strict quality control protocols. However, we stu

Cited by 0SourcePDFScholar
2025

LoCt-Instruct: An Automatic Pipeline for Constructing Datasets of Logical Continuous Instructions

EMNLP 2025

Continuous instruction following closely mirrors real-world tasks by requiring models to solve sequences of interdependent steps, yet existing multi-step instruction datasets suffer from three key limitations: (1) lack of logical coherence across turns, (2) narrow topical breadth and depth, and (3)

2025

Multilingual Dialogue Generation and Localization with Dialogue Act Scripting

EMNLP 2025

Non-English dialogue datasets are scarce, and models are often trained or evaluated on translations of English-language dialogues, an approach which can introduce artifacts that reduce their naturalness and cultural appropriateness. This work proposes Dialogue Act Script (DAS), a structured framewor

Cited by 0SourcePDFScholar
2025

Reliability Crisis of Reference-free Metrics for Grammatical Error Correction

EMNLP 2025

Reference-free evaluation metrics for grammatical error correction (GEC) have achieved high correlation with human judgments.However, these metrics are not designed to evaluate adversarial systems that aim to obtain unjustifiably high scores. The existence of such systems undermines the reliability

2025

Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human?

ACL 2025short

One of the goals of automatic evaluation metrics in grammatical error correction (GEC) is to rank GEC systems such that it matches human preferences. However, current automatic evaluations are based on procedures that diverge from human evaluation. Specifically, human evaluation derives rankings by…

2025

Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability

ACL 2025long

In generative commonsense reasoning tasks such as CommonGen, generative large language models (LLMs) compose sentences that include all given concepts. However, when focusing on instruction-following capabilities, if a prompt specifies a concept order, LLMs must generate sentences that adhere to the…

2025

SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala

EMNLP 2025

Large Language Models (LLMs) demonstrate impressive general knowledge and reasoning abilities, yet their evaluation has predominantly focused on global or anglocentric subjects, often neglecting low-resource languages and culturally specific content. While recent multilingual benchmarks attempt to b

Cited by 0SourcePDFScholar
2025

Tonguescape: Exploring Language Models Understanding of Vowel Articulation

NAACL 2025long

Vowels are primarily characterized by tongue position. Humans have discovered these features of vowel articulation through their own experience and explicit objective observation such as using MRI. With this knowledge and our experience, we can explain and understand the relationship between tongue…

2025

Towards Cross-Lingual Explanation of Artwork in Large-scale Vision Language Models

NAACL 2025findings

As the performance of Large-scale Vision Language Models (LVLMs) improves, they are increasingly capable of responding in multiple languages, and there is an expectation that the demand for explanations generated by LVLMs will grow. However, pre-training of Vision Encoder and the integrated training…

Cited by 5SourcePDFScholar
2024

Centroid-Based Efficient Minimum Bayes Risk Decoding

ACL 2024findings

Minimum Bayes risk (MBR) decoding achieved state-of-the-art translation performance by using COMET, a neural metric that has a high correlation with human evaluation.However, MBR decoding requires quadratic time since it computes the expected score between a translation hypothesis and all reference…

2024

Does Pre-trained Language Model Actually Infer Unseen Links in Knowledge Graph Completion?

NAACL 2024long

Knowledge graphs (KGs) consist of links that describe relationships between entities. Due to the difficulty of manually enumerating all relationships between entities, automatically completing them is essential for KGs. Knowledge Graph Completion (KGC) is a task that infers unseen relationships betw…

Cited by 5SourcePDFScholar
2024

Simul-MuST-C: Simultaneous Multilingual Speech Translation Corpus Using Large Language Model

EMNLP 2024main

Simultaneous Speech Translation (SiST) begins translating before the entire source input is received, making it crucial to balance quality and latency. In real interpreting situations, interpreters manage this simultaneity by breaking sentences into smaller segments and translating them while mainta…

2024

Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair

EMNLP 2024main

In Simultaneous Machine Translation (SiMT), training with a simultaneous interpretation (SI) corpus is an effective method for achieving high-quality yet low-latency. However, constructing such a corpus is challenging due to high costs, and limitations in annotator capabilities, and as a result, exi…

2024

Towards Artwork Explanation in Large-scale Vision Language Models

ACL 2024short

Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating advanced capabilities in text generation and comprehension. However, it has not been clarified to what extent LVLMs understand the knowledge necessary for explaining images, the complex relationships b…

2024

mCSQA: Multilingual Commonsense Reasoning Dataset with Unified Creation Strategy by Language Models and Humans

ACL 2024findings

It is very challenging to curate a dataset for language-specific knowledge and common sense in order to evaluate natural language understanding capabilities of language models. Due to the limitation in the availability of annotators, most current multilingual datasets are created through translation…

Cited by 10SourcePDFScholar
2022

Turning a Functional Cloth Into an Actuator by Combining Thread-Like Thin Artificial Muscles and Embroidery Techniques

RA-L 2022

In this letter, we propose to turn various existing cloths, especially functional cloths, into actuators by combined embroidery technology and thin McKibben artificial muscles. Most of the textile sensors so far are made by weaving active yarns into each other or into other base cloths. These device

Cited by 11SourceScholar
2019

Development and Experimental Validation of Aerial Vehicle With Passive Rotating Shell on Each Rotor

RA-L 2019

Aerial robotics is a fast-growing field of robotics and has been successfully used in various applications. Still, it faces many challenges, such as dealing with unavoidable obstacles in a cluttered environment. Recently, a flying robot with a protective shell that can rotate passively was introduce

Cited by 40SourceScholar