← Search

Jicheol Park

4 accepted papers

2025

Improving Sound Source Localization with Joint Slot Attention on Image and Audio

CVPR 2025poster

Sound source localization (SSL) is the task of locating the source of sound within an image. Due to the lack of localization labels, the de facto standard in SSL has been to represent an image and audio as a single embedding vector each, and use them to learn SSL via contrastive learning. To this en…

Cited by 0SourcePDFScholar
2025

Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval

CVPR 2025poster

Video-text retrieval, the task of retrieving videos based on a textual query or vice versa, is of paramount importance for video understanding and multimodal information retrieval. Recent methods in this area rely primarily on visual and textual features and often ignore audio, although it helps enh…

Cited by 0SourcePDFScholar
2024

PLOT: Text-based Person Search with Part Slot Attention for Corresponding Part Discovery

ECCV 2024poster

"Text-based person search, employing free-form text queries to identify individuals within a vast image collection, presents a unique challenge in aligning visual and textual representations, particularly at the human part level. Existing methods often struggle with part feature extraction and align…

Cited by 4SourcePDFScholar
2021

ASMR: Learning Attribute-Based Person Search With Adaptive Semantic Margin Regularizer

ICCV 2021poster

Attribute-based person search is the task of finding person images that are best matched with a set of text attributes given as query. The main challenge of this task is the large modality gap between attributes and images. To reduce the gap, we present a new loss for learning cross-modal embeddings…

Cited by 29PDFcodeScholar