← Search

Masaaki Nagata

12 accepted papers

2025

Improving Word Alignment Using Semi-Supervised Learning

ACL 2025finding

Word alignment plays a crucial role in various natural language processing tasks, such as serving as cross-lingual signals for sentence embedding, reducing hallucination and omission in machine translation, and facilitating the construction of training data for simultaneous speech translation.Curren…

2024

JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus

COLING 2024main

We constructed JaParaPat (Japanese-English Parallel Patent Application Corpus), a bilingual corpus of more than 300 million Japanese-English sentence pairs from patent applications published in Japan and the United States from 2000 to 2021. We obtained the publication of unexamined patent applicatio…

Cited by 2SourcePDFScholar
2024

Word Alignment as Preference for Machine Translation

EMNLP 2024main

The problem of hallucination and omission, a long-standing problem in machine translation (MT), is more pronounced when a large language model (LLM) is used in MT because an LLM itself is susceptible to these phenomena. In this work, we mitigate the problem in an LLM-based MT model by guiding it to…

2023

WSPAlign: Word Alignment Pre-training via Large-Scale Weakly Supervised Span Prediction

ACL 2023long

Most existing word alignment methods rely on manual alignment datasets or parallel corpora, which limits their usefulness. Here, to mitigate the dependence on manual data, we broaden the source of supervision by relaxing the requirement for correct, fully-aligned, and parallel sentences. Specificall…

2022

A Simple and Strong Baseline for End-to-End Neural RST-style Discourse Parsing

EMNLP 2022finding

To promote and further develop RST-style discourse parsing models, we need a strong baseline that can be regarded as a reference for reporting reliable experimental results. This paper explores a strong baseline by integrating existing simple parsing strategies, top-down and bottom-up, with various…

2022

Domain Adaptation of Machine Translation with Crowdworkers

EMNLP 2022industry

Although a machine translation model trained with a large in-domain parallel corpus achieves remarkable results, it still works poorly when no in-domain data are available. This situation restricts the applicability of machine translation when the target domain’s data are limited. However, there is…

2021

Improving Neural RST Parsing Model with Silver Agreement Subtrees

NAACL 2021long

Most of the previous Rhetorical Structure Theory (RST) parsing methods are based on supervised learning such as neural networks, that require an annotated corpus of sufficient size and quality. However, the RST Discourse Treebank (RST-DT), the benchmark corpus for RST parsing in English, is small du…

2020

SODA: Story Oriented Dense Video Captioning Evaluation Framework

ECCV 2020poster

Dense Video Captioning (DVC) is a challenging task that localizes all events in a short video and describes them with natural language sentences. The main goal of DVC is video story description, that is, to generate a concise video story that supports human video comprehension without watching it. I…

2020

SpanAlign: Sentence Alignment Method based on Cross-Language Span Prediction and ILP

COLING 2020main

We propose a novel method of automatic sentence alignment from noisy parallel documents. We first formalize the sentence alignment problem as the independent predictions of spans in the target document from sentences in the source document. We then introduce a total optimization method using integer…

2019

ILP-based Compressive Speech Summarization with Content Word Coverage Maximization and Its Oracle Performance Analysis

ICASSP 2019accepted

We propose an integer linear programming (ILP)-based compressive speech summarization method that maximizes the coverage of content words in a resultant summary. It is an unsupervised method and, under the designed constraints, it performs a single-step globally optimal summarization of a given long…

Cited by 0SourceScholar