← Search

Bhavin Jawade

2 accepted papers

2025

Audio-Visual Representation Learning For Lip-Sync Estimation Through Ranking Augmented Contrastive Training

ICASSP 2025accepted

In many applications, particularly in media production and content localization, it is crucial to detect and evaluate varying degrees of audio-visual synchronization, such as selecting high-quality dubbed audio over poorly synchronized tracks. Traditional contrastively pre-trained LipSync models are…

Cited by 0SourceScholar
2024

ProxyFusion: Face Feature Aggregation Through Sparse Experts

NeurIPS 2024poster

Face feature fusion is indispensable for robust face recognition, particularly in scenarios involving long-range, low-resolution media (unconstrained environments) where not all frames or features are equally informative. Existing methods often rely on large intermediate feature maps or face metadat…