← Search

Hiroya Takamura

31 accepted papers

2026

Low-Latency Real-Time Audio Game Commentary System via LLM-based Parallel Text Generation

IJCAI 2026

We present a low-latency real-time audio game commentary system that generates spoken commentary directly from live gameplay video. In this end-to-end setting, a key bottleneck is accumulated waiting time; conventional pipelines capture frames, generate text, and synthesize speech sequentially for e

Cited by 0Scholar
2026

Rewriting Pre-Training Data Boosts LLM Performance in Math and Code

ICLR 2026poster

The performance of large language models (LLMs) in program synthesis and mathematical reasoning is fundamentally limited by the quality of their pre-training corpora. We introduce two openly licensed pre-training datasets, released under the Llama 3.3 Community License, that significantly enhance…

Cited by 0SourcecodeScholar
2025

AdParaphrase v2.0: Generating Attractive Ad Texts Using a Preference-Annotated Paraphrase Dataset

ACL 2025finding

Identifying factors that make ad text attractive is essential for advertising success. This study proposes AdParaphrase v2.0, a dataset for ad text paraphrasing, containing human preference data, to enable the analysis of the linguistic factors and to support the development of methods for generatin…

2025

AdParaphrase: Paraphrase Dataset for Analyzing Linguistic Features toward Generating Attractive Ad Texts

NAACL 2025findings

Effective linguistic choices that attract potential customers play crucial roles in advertising success. This study aims to explore the linguistic features of ad texts that influence human preferences. Although the creation of attractive ad texts is an active area of research, progress in understand…

2025

Analyzing Continuous Semantic Shifts with Diachronic Word Similarity Matrices

COLING 2025main

The meanings and relationships of words shift over time. This phenomenon is referred to as semantic shift. Research focused on understanding how semantic shifts occur over multiple time periods is essential for gaining a detailed understanding of semantic shifts. However, detecting change points onl…

2025

Can GPT-4 Sway Experts’ Investment Decisions?

NAACL 2025findings

In the post-Turing era, evaluating large language models (LLMs) involves assessing generated text based on readers’ decisions rather than merely its indistinguishability from human-produced content. This paper explores how LLM-generated text impacts readers’ decisions, focusing on both amateur and e…

2025

GADFA: Generator-Assisted Decision-Focused Approach for Opinion Expressing Timing Identification

COLING 2025main

The advancement of text generation models has granted us the capability to produce coherent and convincing text on demand. Yet, in real-life circumstances, individuals do not continuously generate text or voice their opinions. For instance, consumers pen product reviews after weighing the merits and…

Cited by 0SourcePDFScholar
2024

A Computational Approach to Quantifying Grammaticization of English Deverbal Prepositions

COLING 2024main

This paper explores grammaticization of deverbal prepositions by a computational approach based on corpus data. Deverbal prepositions are words or phrases that are derived from a verb and that behave as a preposition such as “regarding” and “according to”. Linguistic studies have revealed important…

Cited by 0SourcePDFScholar
2024

DBQR-QA: A Question Answering Dataset on a Hybrid of Database Querying and Reasoning

ACL 2024findings

This paper introduces the Database Querying and Reasoning Dataset for Question Answering (DBQR-QA), aimed at addressing the gap in current question-answering (QA) research by emphasizing the essential processes of database querying and reasoning to answer questions. Specifically designed to accommod…

Cited by 0SourcePDFScholar
2024

Demonstration Selection Strategies for Numerical Time Series Data-to-Text

EMNLP 2024finding

Demonstration selection, the process of selecting examples used in prompts, plays a critical role in in-context learning. This paper explores demonstration selection methods for data-to-text tasks that involve numerical time series data as inputs.Previously developed demonstration selection methods…

Cited by 0SourcePDFScholar
2024

Introducing Spatial Information and a Novel Evaluation Scheme for Open-Domain Live Commentary Generation

EMNLP 2024finding

This paper focuses on the task of open-domain live commentary generation. Compared to domain-specific work in this task, this setting proved particularly challenging due to the absence of domain-specific features. Aiming to bridge this gap, we integrate spatial information by proposing an utterance…

Cited by 1SourcePDFScholar
2024

Prompting for Numerical Sequences: A Case Study on Market Comment Generation

COLING 2024main

Large language models (LLMs) have been applied to a wide range of data-to-text generation tasks, including tables, graphs, and time-series numerical data-to-text settings. While research on generating prompts for structured data such as tables and graphs is gaining momentum, in-depth investigations…

2024

The Impact of Language on Arithmetic Proficiency: A Multilingual Investigation with Cross-Agent Checking Computation

NAACL 2024short

This paper critically examines the arithmetic capabilities of Large Language Models (LLMs), uncovering significant limitations in their performance. Our research reveals a notable decline in accuracy for complex calculations involving large numbers, with addition and subtraction tasks showing varyin…

Cited by 1SourcePDFScholar
2024

Who Said What: Formalization and Benchmarks for the Task of Quote Attribution

COLING 2024main

The task of quote attribution seeks to pair textual utterances with the name of their speakers. Despite continuing research efforts on the task, models are rarely evaluated systematically against previous models in comparable settings on the same datasets. This has resulted in a poor understanding o…

2023

Towards Parameter-Efficient Integration of Pre-Trained Language Models In Temporal Video Grounding

ACL 2023findings

This paper explores the task of Temporal Video Grounding (TVG) where, given an untrimmed video and a query sentence, the goal is to recognize and determine temporal boundaries of action instances in the video described by natural language queries. Recent works tackled this task by improving query in…

2023

Variance Matters: Detecting Semantic Differences without Corpus/Word Alignment

EMNLP 2023long main

In this paper, we propose methods for discovering semantic differences in words appearing in two corpora. The key idea is to measure the coverage of meanings of a word in a corpus through the norm of its mean word vector, which is equivalent to examining a kind of variance of the word vector distrib…

Cited by 0SourceScholar
2022

Aspect-based Analysis of Advertising Appeals for Search Engine Advertising

NAACL 2022industry

Writing an ad text that attracts people and persuades them to click or act is essential for the success of search engine advertising. Therefore, ad creators must consider various aspects of advertising appeals (A3) such as the price, product features, and quality. However, products and services exhi…

Cited by 3SourcePDFScholar
2022

Infinite SCAN: An Infinite Model of Diachronic Semantic Change

EMNLP 2022main

In this study, we propose a Bayesian model that can jointly estimate the number of senses of words and their changes through time.The model combines a dynamic topic model on Gaussian Markov random fields with a logistic stick-breaking process that realizes Dirichlet process. In the experiments, we e…

Cited by 6SourcePDFScholar
2022

Open-domain Video Commentary Generation

EMNLP 2022main

Live commentary plays an important role in sports broadcasts and video games, making spectators more excited and immersed. In this context, though approaches for automatically generating such commentary have been proposed in the past, they have been generally concerned with specific fields, where it…

2022

Revisiting Statistical Laws of Semantic Shift in Romance Cognates

COLING 2022main

This article revisits statistical relationships across Romance cognates between lexical semantic shift and six intra-linguistic variables, such as frequency and polysemy. Cognates are words that are derived from a common etymon, in this case, a Latin ancestor. Despite their shared etymology, some co…

Cited by 4SourcePDFScholar
2022

StoryER: Automatic Story Evaluation via Ranking, Rating and Reasoning

EMNLP 2022main

Existing automatic story evaluation methods place a premium on story lexical level coherence, deviating from human preference.We go beyond this limitation by considering a novel Story Evaluation method that mimics human preference when judging a story, namely StoryER, which consists of three sub-tas…

2021

An Empirical Study of Generating Texts for Search Engine Advertising

NAACL 2021industry

Although there are many studies on neural language generation (NLG), few trials are put into the real world, especially in the advertising domain. Generating ads with NLG models can help copywriters in their creation. However, few studies have adequately evaluated the effect of generated ads with ac…

2021

Commonsense Knowledge Aware Concept Selection For Diverse and Informative Visual Storytelling

AAAI 2021technical

Visual storytelling is a task of generating relevant and interesting stories for given image sequences. In this work we aim at increasing the diversity of the generated stories while preserving the informative content from the images. We propose to foster the diversity and informativeness of a gener…

Cited by 49SourcePDFScholar
2021

SciXGen: A Scientific Paper Dataset for Context-Aware Text Generation

EMNLP 2021finding

Generating texts in scientific papers requires not only capturing the content contained within the given input but also frequently acquiring the external information called context. We push forward the scientific text generation by proposing a new task, namely context-aware text generation in the sc…

Cited by 20SourcePDFScholar
2021

Towards Table-to-Text Generation with Numerical Reasoning

ACL 2021long

Recent neural text generation models have shown significant improvement in generating descriptive text from structured data such as table formats. One of the remaining important challenges is generating more analytical descriptions that can be inferred from facts in a data source. The use of a templ…

2020

A Neural Model for Aggregating Coreference Annotation in Crowdsourcing

COLING 2020main

Coreference resolution is the task of identifying all mentions in a text that refer to the same real-world entity. Collecting sufficient labelled data from expert annotators to train a high-performance coreference resolution system is time-consuming and expensive. Crowdsourcing makes it possible to…

Cited by 11SourcePDFScholar
2020

An empirical analysis of existing systems and datasets toward general simple question answering

COLING 2020main

In this paper, we evaluate the progress of our field toward solving simple factoid questions over a knowledge base, a practically important problem in natural language interface to database. As in other natural language understanding tasks, a common practice for this task is to train and evaluate a…

2020

Learning with Contrastive Examples for Data-to-Text Generation

COLING 2020main

Existing models for data-to-text tasks generate fluent but sometimes incorrect sentences e.g., “Nikkei gains” is generated when “Nikkei drops” is expected. We investigate models trained on contrastive examples i.e., incorrect sentences or terms, in addition to correct ones to reduce such errors. We…

2020

Neural text normalization leveraging similarities of strings and sounds

COLING 2020main

We propose neural models that can normalize text by considering the similarities of word strings and sounds. We experimentally compared a model that considers the similarities of both word strings and sounds, a model that considers only the similarity of word strings or of sounds, and a model withou…

Cited by 5SourcePDFScholar
2020

Pointing to Subwords for Generating Function Names in Source Code

COLING 2020main

We tackle the task of automatically generating a function name from source code. Existing generators face difficulties in generating low-frequency or out-of-vocabulary subwords. In this paper, we propose two strategies for copying low-frequency or out-of-vocabulary subwords in inputs. Our best perfo…