← Search

Shota Nakada

3 accepted papers

2025

DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information

ICASSP 2025accepted

Current audio-visual representation learning can capture rough object categories (e.g., "animals" and "instruments"), but it lacks the ability to recognize fine-grained details, such as specific categories like "dogs" and "flutes" within animals and instruments. To address this issue, we introduce D…

Cited by 0SourceScholar
2024

Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection

EMNLP 2024system demonstrations

We propose Lighthouse, a user-friendly library for reproducible video moment retrieval and highlight detection (MR-HD). Although researchers proposed various MR-HD approaches, the research community holds two main issues. The first is a lack of comprehensive and reproducible experiments across vario…