← Search

Carlos-D. Martínez-Hinarejos

2 accepted papers

2024

AnnoTheia: A Semi-Automatic Annotation Toolkit for Audio-Visual Speech Technologies

COLING 2024main

More than 7,000 known languages are spoken around the world. However, due to the lack of annotated resources, only a small fraction of them are currently covered by speech technologies. Albeit self-supervised speech representations, recent massive speech corpora collections, as well as the organizat…

2024

Comparison of Conventional Hybrid and CTC/Attention Decoders for Continuous Visual Speech Recognition

COLING 2024main

Thanks to the rise of deep learning and the availability of large-scale audio-visual databases, recent advances have been achieved in Visual Speech Recognition (VSR). Similar to other speech processing tasks, these end-to-end VSR systems are usually based on encoder-decoder architectures. While enco…