Fusing Information Streams in End-to-End Audio-Visual Speech Recognition
End-to-end acoustic speech recognition has quickly gained widespread popularity and shows promising results in many studies. Specifically the joint transformer/CTC model pro-vides very good performance in many tasks. However, under noisy and distorted conditions, the performance still degrades notab…