2025
Audio-Visual Representation Learning For Lip-Sync Estimation Through Ranking Augmented Contrastive Training
ICASSP 2025accepted
In many applications, particularly in media production and content localization, it is crucial to detect and evaluate varying degrees of audio-visual synchronization, such as selecting high-quality dubbed audio over poorly synchronized tracks. Traditional contrastively pre-trained LipSync models are…