LinProVSR: Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition
Feng Xue, Baochao Zhu, Wei Jia, Shujie Li, Yu Li, Jinrui Zhang, Shengeng Tang, Dan Guo
Abstract
Visual Speech Recognition (VSR), commonly known as lipreading, enables the recognition of spoken text by analyzing lip visual features. Due to the subtlety of lip movements, its recognition is much harder than other motion recognition tasks. Existing VSR models face the challenge of viseme ambiguity when processing phonemes with similar pronunciations—multiple phonemes share similar viseme features, leading to a notable drop in lipreading accuracy. To address this issue, this study proposes a Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition(LinProVSR) framework. First, an ambiguous sample set is constructed based on linguistic knowledge to provide supervisory signals for the model
BibTeX
@inproceedings{aaai2026_linprovsrlinguis,
title = {LinProVSR: Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition},
author = {Feng Xue and Baochao Zhu and Wei Jia and Shujie Li and Yu Li and Jinrui Zhang and Shengeng Tang and Dan Guo},
booktitle = {AAAI 2026},
year = {2026}
}