2025
DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
ICASSP 2025accepted
Current audio-visual representation learning can capture rough object categories (e.g., "animals" and "instruments"), but it lacks the ability to recognize fine-grained details, such as specific categories like "dogs" and "flutes" within animals and instruments. To address this issue, we introduce D…