Neural-Inspired Modeling of Auditory Selection and Compensation for Audio-Visual Speech Separation
Current audio-visual speech separation (AVSS) models typically rely on implicit multimodal fusion, but the absence of explicit modality alignment and reliability modeling often causes semantic misalignment and contaminates speech representations. The brain addresses this with a hierarchy: top-down a…