← Search

Chihiro Taguchi

6 accepted papers

2025

Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive‐k

EMNLP 2025

Retrieval-augmented generation (RAG) and long-context language models (LCLMs) both address context limitations of LLMs in open-domain QA. However, optimal external context to retrieve remains an open problem: fixed retrieval budgets risk wasting tokens or omitting key evidence. Existing adaptive met

Cited by 0SourcePDFScholar
2025

Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark

EMNLP 2025

Multilingual machine translation (MT) benchmarks play a central role in evaluating the capabilities of modern MT systems. Among them, the FLORES+ benchmark is widely used, offering English-to-many translation data for over 200 languages, curated with strict quality control protocols. However, we stu

Cited by 0SourcePDFScholar
2025

SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches

ICLR 2025poster

Researchers and practitioners in natural language processing and computational linguistics frequently observe and analyze the real language usage in large-scale corpora. For that purpose, they often employ off-the-shelf pattern-matching tools, such as grep, and keyword-in-context concordancers, whic…

Cited by 0SourcePDFScholar
2024

J-SNACS: Adposition and Case Supersenses for Japanese Joshi

COLING 2024main

Many languages use adpositions (prepositions or postpositions) to mark a variety of semantic relations, with different languages exhibiting both commonalities and idiosyncrasies in the relations grouped under the same lexeme. We present the first Japanese extension of the SNACS framework (Schneider…

2024

Killkan: The Automatic Speech Recognition Dataset for Kichwa with Morphosyntactic Information

COLING 2024main

This paper presents Killkan, the first dataset for automatic speech recognition (ASR) in the Kichwa language, an indigenous language of Ecuador. Kichwa is an extremely low-resource endangered language, and there have been no resources before Killkan for Kichwa to be incorporated in applications of n…

2024

Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn’t

ACL 2024long

We investigate what linguistic factors affect the performance of Automatic Speech Recognition (ASR) models. We hypothesize that orthographic and phonological complexities both degrade accuracy. To examine this, we fine-tune the multilingual self-supervised pretrained model Wav2Vec2-XLSR-53 on 25 lan…