The Potential of Speech Features to Discriminate between Original and Machine-Translated Texts
Yongjian Chen, Mireia Farrús, Antonio Toral
Abstract
Discriminating between original texts and machine translations involves identifying whether a text was originally authored in the target language or generated through machine translation. To our knowledge, all methods to date depend exclusively on text-based features. In this study, we move beyond this unimodal approach by incorporating speech features. Machine-translated texts display linguistic deviations from original texts, such as those in lexicon and syntax, which can also manifest in speech characteristics. We evaluate the effectiveness of using text features, speech features, and their bimodal fusion to train classifiers capable of discerning original from machine-translated texts. Additionally, we explore various classification algorithms and fusion techniques. Our results show that speech features alone surpass chance accuracy, while combining text and speech features enhances performance beyond text-only methods. Furthermore, although no single classification or fusion method proves consistently superior, advanced fusion techniques outperform simple feature concatenation.
BibTeX
@inproceedings{icassp2025_thepotentialofsp,
title = {The Potential of Speech Features to Discriminate between Original and Machine-Translated Texts},
author = {Yongjian Chen and Mireia Farrús and Antonio Toral},
booktitle = {ICASSP 2025},
year = {2025}
}