← Search

Hiroki Ouchi

12 accepted papers

2026

VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction

AAAI 2026technical

Recent advances in multimodal large language models (MLLMs) have significantly enhanced video understanding capabilities, opening new possibilities for practical applications. Yet current video benchmarks focus largely on indoor scenes or short-range outdoor activities, leaving the challenges associ

Cited by 0SourcePDFScholar
2025

A Text Embedding Model with Contrastive Example Mining for Point-of-Interest Geocoding

COLING 2025main

Geocoding is a fundamental technique that links location mentions to their geographic positions, which is important for understanding texts in terms of where the described events occurred. Unlike most geocoding studies that targeted coarse-grained locations, we focus on geocoding at a fine-grained p…

2025

AdTEC: A Unified Benchmark for Evaluating Text Quality in Search Engine Advertising

NAACL 2025long

As the fluency of ad texts automatically generated by natural language generation technologies continues to improve, there is an increasing demand to assess the quality of these creatives in real-world setting.We propose **AdTEC**, the first public benchmark to evaluate ad texts from multiple perspe…

2025

BannerBench: Benchmarking Vision Language Models for Multi-Ad Selection with Human Preferences

EMNLP 2025

Web banner advertisements, which are placed on websites to guide users to a targeted landing page (LP), are still often selected manually because human preferences are important in selecting which ads to deliver. To automate this process, we propose a new benchmark, BannerBench, to evaluate the huma

Cited by 0SourcePDFScholar
2025

Graph-Structured Trajectory Extraction from Travelogues

ACL 2025long

Human traveling trajectories play a central role in characterizing each travelogue, and automatic trajectory extraction from travelogues is highly desired for tourism services, such as travel planning and recommendation. This work addresses the extraction of human traveling trajectories from travelo…

2024

Can Language Models Induce Grammatical Knowledge from Indirect Evidence?

EMNLP 2024main

What kinds of and how much data is necessary for language models to induce grammatical knowledge to judge sentence acceptability? Recent language models still have much room for improvement in their data efficiency compared to humans. This paper investigates whether language models efficiently use i…

2024

Modeling Overregularization in Children with Small Language Models

ACL 2024findings

The imitation of the children’s language acquisition process has been explored to make language models (LMs) more efficient.In particular, errors caused by children’s regularization (so-called overregularization, e.g., using wroted for the past tense of write) have been widely studied to reveal the…

2023

Second Language Acquisition of Neural Language Models

ACL 2023findings

With the success of neural language models (LMs), their language acquisition has gained much attention. This work sheds light on the second language (L2) acquisition of LMs, while previous work has typically explored their first language (L1) acquisition. Specifically, we trained bilingual LMs with…

2022

Iterative Span Selection: Self-Emergence of Resolving Orders in Semantic Role Labeling

COLING 2022main

Semantic Role Labeling (SRL) is the task of labeling semantic arguments for marked semantic predicates. Semantic arguments and their predicates are related in various distinct manners, of which certain semantic arguments are a necessity while others serve as an auxiliary to their predicates. To cons…

Cited by 2SourcePDFScholar
2021

Pseudo Zero Pronoun Resolution Improves Zero Anaphora Resolution

EMNLP 2021main

Masked language models (MLMs) have contributed to drastic performance improvements with regard to zero anaphora resolution (ZAR). To further improve this approach, in this study, we made two proposals. The first is a new pretraining task that trains MLMs on anaphoric relations with explicit supervis…

2020

An Empirical Study of Contextual Data Augmentation for Japanese Zero Anaphora Resolution

COLING 2020main

One critical issue of zero anaphora resolution (ZAR) is the scarcity of labeled data. This study explores how effectively this problem can be alleviated by data augmentation. We adopt a state-of-the-art data augmentation method, called the contextual data augmentation (CDA), that generates labeled t…

Cited by 9SourcePDFScholar