← Search

Jinyoung Moon

7 accepted papers

2025

HiCM²: Hierarchical Compact Memory Modeling for Dense Video Captioning

AAAI 2025technical

With the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and localization of untrimmed videos. Several studies highlight the challenges of DVC and introduce improved methods utilizing pr…

Cited by 1SourcePDFScholar
2025

Weakly Supervised Video Scene Graph Generation via Natural Language Supervision

ICLR 2025poster

Existing Video Scene Graph Generation (VidSGG) studies are trained in a fully supervised manner, which requires all frames in a video to be annotated, thereby incurring high annotation cost compared to Image Scene Graph Generation (ImgSGG). Although the annotation cost of VidSGG can be alleviated by…

2024

Adaptive Self-training Framework for Fine-grained Scene Graph Generation

ICLR 2024poster

Scene graph generation (SGG) models have suffered from inherent problems regarding the benchmark datasets such as the long-tailed predicate distribution and missing annotation problems. In this work, we aim to alleviate the long-tailed problem of SGG by utilizing unannotated triplets. To this end, w…

2024

Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval

CVPR 2024poster

There has been significant attention to the research on dense video captioning which aims to automatically localize and caption all events within untrimmed video. Several studies introduce methods by designing dense video captioning as a multitasking problem of event localization and event captionin…

2024

LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation

CVPR 2024poster

Weakly-Supervised Scene Graph Generation (WSSGG) research has recently emerged as an alternative to the fully-supervised approach that heavily relies on costly annotations. In this regard studies on WSSGG have utilized image captions to obtain unlocalized triplets while primarily focusing on groundi…

2023

Unbiased Heterogeneous Scene Graph Generation with Relation-Aware Message Passing Neural Network

AAAI 2023technical

Recent scene graph generation (SGG) frameworks have focused on learning complex relationships among multiple objects in an image. Thanks to the nature of the message passing neural network (MPNN) that models high-order interactions between objects and their neighboring objects, they are dominant rep…

2020

Learning to Discriminate Information for Online Action Detection

CVPR 2020poster

From a streaming video, online action detection aims to identify actions in the present. For this task, previous methods use recurrent networks to model the temporal sequence of current action frames. However, these methods overlook the fact that an input image sequence includes background and irrel…

Cited by 100PDFScholar