← Search

Junyeong Kim

12 accepted papers

2026

GranAlign: Granularity-Aware Alignment Framework for Zero-shot Video Moment Retrieval

AAAI 2026technical

Zero-shot video moment retrieval (ZVMR) is the task of localizing a temporal moment within an untrimmed video using a natural language query without relying on task-specific training data. The primary challenge in this setting lies in the mismatch in semantic granularity between textual queries and

Cited by 0SourcePDFScholar
2025

Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio Captioning

EMNLP 2025

Automated Audio Captioning (AAC) aims to generate natural language descriptions of audio content, enabling machines to interpret and communicate complex acoustic scenes. However, current AAC datasets often suffer from short and simplistic captions, limiting model expressiveness and semantic depth. T

Cited by 0SourcePDFScholar
2025

QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering

EMNLP 2025

Video-to-text summarization remains underexplored in terms of comprehensive evaluation methods. Traditional n-gram overlap-based metrics and recent large language model (LLM)-based approaches depend heavily on human-written reference summaries, limiting their practicality and sensitivity to nuanced

2023

Counterfactual Two-Stage Debiasing For Video Corpus Moment Retrieval

ICASSP 2023accepted

Video Corpus Moment Retrieval aims to select a temporal video moment pertinent to a given language query from a large video corpus. Existing systems are prone to rely on a retrieval bias as a shortcut, which hinders the systems from accurately learning vision-language association. The retrieval bias…

Cited by 0SourceScholar
2023

HEAR: Hearing Enhanced Audio Response for Video-grounded Dialogue

EMNLP 2023long findings

Video-grounded Dialogue (VGD) aims to answer questions regarding a given multi-modal input comprising video, audio, and dialogue history. Although there have been numerous efforts in developing VGD systems to improve the quality of their responses, existing systems are competent only to incorporate…

Cited by 0SourcecodeScholar
2022

Information-Theoretic Text Hallucination Reduction for Video-grounded Dialogue

EMNLP 2022main

Video-grounded Dialogue (VGD) aims to decode an answer sentence to a question regarding a given video and dialogue context. Despite the recent success of multi-modal reasoning to generate answer sentences, existing dialogue systems still suffer from a text hallucination problem, which denotes indisc…

2022

Selective Query-Guided Debiasing for Video Corpus Moment Retrieval

ECCV 2022poster

"Video moment retrieval (VMR) aims to localize target moments in untrimmed videos pertinent to a given textual query. Existing retrieval systems tend to rely on retrieval bias as a shortcut and thus, fail to sufficiently learn multi-modal interactions between query and video. This retrieval bias ste…

2021

Structured Co-reference Graph Attention for Video-grounded Dialogue

AAAI 2021technical

A video-grounded dialogue system referred to as the Structured Co-reference Graph Attention (SCGA) is presented for decoding the answer sequence to a question regarding a given video while keeping track of the dialogue context. Although recent efforts have made great strides in improving the quality…

Cited by 26SourcePDFScholar
2020

Modality Shifting Attention Network for Multi-Modal Video Question Answering

CVPR 2020poster

This paper considers a network referred to as Modality Shifting Attention Network (MSAN) for Multimodal Video Question Answering (MVQA) task. MSAN decomposes the task into two sub-tasks: (1) localization of temporal moment relevant to the question, and (2) accurate prediction of the answer based on…

Cited by 102PDFScholar
2020

VLANet: Video-Language Alignment Network for Weakly-Supervised Video Moment Retrieval

ECCV 2020poster

Video Moment Retrieval (VMR) is a task to localize the temporal moment in untrimmed video specified by natural language query. For VMR, several methods that require full supervision for training have been proposed. Unfortunately, acquiring a large number of training videos with labeled temporal boun…

Cited by 99SourcePDFScholar
2019

Progressive Attention Memory Network for Movie Story Question Answering

CVPR 2019poster

This paper proposes the progressive attention memory network (PAMN) for movie story question answering (QA). Movie story QA is challenging compared to VQA in two aspects: (1) pinpointing the temporal parts relevant to answer the question is difficult as the movies are typically longer than an hour,…

Cited by 96PDFScholar
2018

Pivot Correlational Neural Network for Multimodal Video Categorization

ECCV 2018poster

This paper considers an architecture for multimodal video categorization referred to as Pivot Correlational Neural Network (Pivot CorrNN). The architecture is trained to maximizes the correlation between the hidden states as well as the predictions of the modal-agnostic pivot stream and modal-specif…

Cited by 14SourcePDFScholar