← Search

Jonghee Kim

3 accepted papers

2026

GranAlign: Granularity-Aware Alignment Framework for Zero-shot Video Moment Retrieval

AAAI 2026technical

Zero-shot video moment retrieval (ZVMR) is the task of localizing a temporal moment within an untrimmed video using a natural language query without relying on task-specific training data. The primary challenge in this setting lies in the mismatch in semantic granularity between textual queries and

Cited by 0SourcePDFScholar
2023

Exploring The Role of Mean Teachers in Self-supervised Masked Auto-Encoders

ICLR 2023poster

Masked image modeling (MIM) has become a popular strategy for self-supervised learning (SSL) of visual representations with Vision Transformers. A representative MIM model, the masked auto-encoder (MAE), randomly masks a subset of image patches and reconstructs the masked patches given the unmasked…

2022

MPViT: Multi-Path Vision Transformer for Dense Prediction

CVPR 2022poster

Dense computer vision tasks such as object detection and segmentation require effective multi-scale feature representation for detecting or classifying objects or regions with varying sizes. While Convolutional Neural Networks (CNNs) have been the dominant architectures for such tasks, recently intr…

Cited by 360PDFcodeScholar