ICASSP 2025accepted0 citations

Semi-Supervised Multilingual Alignment with Lexical Memory for Massively Parallel Text Mining

Weitai Zhang, Peiwang Tang, Chao Lin, Simran Naagar, Zhongyi Ye, Junhua Liu

Abstract

Existing state-of-the-art techniques that employ multilingual sentence embeddings for mining parallel texts predominantly rely on extensive supervision, which often results in sub-optimal performance in the absence of large-scale parallel training datasets. In this study, we introduce a novel method designed to extract high-quality parallel texts from monolingual corpora, particularly targeting zero-and low-resource languages. We learn language-agnostic sentence embeddings with a two-tiered training regimen: an initial phase of lexical knowledge-enhanced pretraining and a subsequent phase of supervised fine-tuning on a minimally sized parallel dataset using contrastive loss. Furthermore, we enhance the model’s performance by adopting an iterative training methodology that leverages both mined data and synthetically augmented data. We illustrate the capability of our method to create high-quality parallel text with various downstream tasks. All results suggest that the proposed method is effective and can surpass previous state-of-the-art supervised methods in zero-and low-resource scenarios.

BibTeX
@inproceedings{icassp2025_semisupervisedmu,
  title = {Semi-Supervised Multilingual Alignment with Lexical Memory for Massively Parallel Text Mining},
  author = {Weitai Zhang and Peiwang Tang and Chao Lin and Simran Naagar and Zhongyi Ye and Junhua Liu},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Semi-Supervised Multilingual Alignment with Lexical Memory for Massively Parallel Text Mining · ICASSP 2025