← Search

Masato Mita

10 accepted papers

2025

AdTEC: A Unified Benchmark for Evaluating Text Quality in Search Engine Advertising

NAACL 2025long

As the fluency of ad texts automatically generated by natural language generation technologies continues to improve, there is an increasing demand to assess the quality of these creatives in real-world setting.We propose **AdTEC**, the first public benchmark to evaluate ad texts from multiple perspe…

2025

BannerBench: Benchmarking Vision Language Models for Multi-Ad Selection with Human Preferences

EMNLP 2025

Web banner advertisements, which are placed on websites to guide users to a targeted landing page (LP), are still often selected manually because human preferences are important in selecting which ads to deliver. To automate this process, we propose a new benchmark, BannerBench, to evaluate the huma

Cited by 0SourcePDFScholar
2025

Developmentally-plausible Working Memory Shapes a Critical Period for Language Acquisition

ACL 2025long

Large language models possess general linguistic abilities but acquire language less efficiently than humans. This study proposes a method for integrating the developmental characteristics of working memory during the critical period, a stage when human language acquisition is particularly efficient…

2025

Targeted Syntactic Evaluation for Grammatical Error Correction

ACL 2025long

Language learners encounter a wide range of grammar items across the beginner, intermediate, and advanced levels.To develop grammatical error correction (GEC) models effectively, it is crucial to identify which grammar items are easier or more challenging for models to correct. However, conventional…

2024

CAMERA³: An Evaluation Dataset for Controllable Ad Text Generation in Japanese

COLING 2024main

Ad text generation is the task of creating compelling text from an advertising asset that describes products or services, such as a landing page. In advertising, diversity plays an important role in enhancing the effectiveness of an ad text, mitigating a phenomenon called “ad fatigue,” where users b…

Cited by 2SourcePDFScholar
2024

Striking Gold in Advertising: Standardization and Exploration of Ad Text Generation

ACL 2024long

In response to the limitations of manual ad creation, significant research has been conducted in the field of automatic ad text generation (ATG). However, the lack of comprehensive benchmarks and well-defined problem sets has made comparing different methods challenging. To tackle these challenges,…

2020

PheMT: A Phenomenon-wise Dataset for Machine Translation Robustness on User-Generated Contents

COLING 2020main

Neural Machine Translation (NMT) has shown drastic improvement in its quality when translating clean input, such as text from the news domain. However, existing studies suggest that NMT still struggles with certain kinds of input with considerable noise, such as User-Generated Contents (UGC) on the…

2020

Taking the Correction Difficulty into Account in Grammatical Error Correction Evaluation

COLING 2020main

This paper presents performance measures for grammatical error correction which take into account the difficulty of error correction. To the best of our knowledge, no conventional measure has such functionality despite the fact that some errors are easy to correct and others are not. The main purpos…