← Search

Ankit Shahy

1 accepted papers

2024

Conformer is All You Need for Visual Speech Recognition

ICASSP 2024accepted

Visual speech recognition models extract visual features in a hierarchical manner. At the lower level, there is a visual front-end with a limited temporal receptive field that processes the raw pixels depicting the lips or faces. At the higher level, there is an encoder that attends to the embedding…

Cited by 0SourceScholar