← Search

Hokuto Munakata

5 accepted papers

2026

CASTELLA: LONG AUDIO DATASET WITH CAPTIONS AND TEMPORAL BOUNDARIES

ICASSP 2026poster

We introduce CASTELLA, a human-annotated audio benchmark for the task of audio moment retrieval (AMR). Although AMR has various useful potential applications, there is still no established benchmark with real-world data. The initial study of AMR trained the models solely on synthetic datasets. Moreo…

Cited by 0SourcePDFScholar
2025

Aligned Contrastive Learning for Text-to-Music Retrieval

ICASSP 2025accepted

This paper proposes aligned contrastive learning for text-to-music retrieval. The proposed method introduces a new similarity measure, 'aligned similarity', which captures the frame-level and token-level correspondence within text and audio sequences. Unlike traditional approaches that aggregate seq…

Cited by 0SourceScholar
2025

DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information

ICASSP 2025accepted

Current audio-visual representation learning can capture rough object categories (e.g., "animals" and "instruments"), but it lacks the ability to recognize fine-grained details, such as specific categories like "dogs" and "flutes" within animals and instruments. To address this issue, we introduce D…

Cited by 0SourceScholar
2024

Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection

EMNLP 2024system demonstrations

We propose Lighthouse, a user-friendly library for reproducible video moment retrieval and highlight detection (MR-HD). Although researchers proposed various MR-HD approaches, the research community holds two main issues. The first is a lack of comprehensive and reproducible experiments across vario…