← Search

Takaki Makino

1 accepted papers

2020

End-to-End Multi-Person Audio/Visual Automatic Speech Recognition

ICASSP 2020accepted

Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However, in a more realistic setting, when multiple faces are potentially on screen one needs to decide which face to feed to the…

Cited by 0SourceScholar