← Search

Eun-Sol Kim

11 accepted papers

2025

Zero-Shot Compositional Video Learning with Coding Rate Reduction

ICCV 2025poster

In this paper, we propose a novel zero-shot compositional video understanding method inspired by how young children efficiently learn new concepts and flexibly expand their existing knowledge framework. While recent large-scale visual language models (VLMs) have achieved remarkable advancements and…

2024

Compositional Video Understanding with Spatiotemporal Structure-based Transformers

CVPR 2024poster

In this paper we suggest a new novel method to understand complex semantic structures through long video inputs. Conventional methods for understanding videos have been focused on short-term clips and trained to get visual representations for the short clips using convolutional neural networks or tr…

2024

Structure-Aware Multimodal Sequential Learning for Visual Dialog

AAAI 2024technical

With the ability to collect vast amounts of image and natural language data from the web, there has been a remarkable advancement in Large-scale Language Models (LLMs). This progress has led to the emergence of chatbots and dialogue systems capable of fluent conversations with humans. As the variety…

Cited by 1SourcePDFScholar
2022

Hypergraph Transformer: Weakly-Supervised Multi-hop Reasoning for Knowledge-based Visual Question Answering

ACL 2022long

Knowledge-based visual question answering (QA) aims to answer a question which requires visually-grounded external knowledge beyond image content itself. Answering complex questions that require multi-hop reasoning under weak supervision is considered as a challenging problem since i) no supervision…

2022

MSTR: Multi-Scale Transformer for End-to-End Human-Object Interaction Detection

CVPR 2022poster

Human-Object Interaction (HOI) detection is the task of identifying a set of <human, object, interaction> triplets from an image. Recent work proposed transformer encoder-decoder architectures that successfully eliminated the need for many hand-designed components in HOI detection through end-to-end…

Cited by 85PDFcodeScholar
2022

Selective Token Generation for Few-shot Natural Language Generation

COLING 2022main

Natural language modeling with limited training data is a challenging problem, and many algorithms make use of large-scale pretrained language models (PLMs) for this due to its great generalization ability. Among them, additive learning that incorporates a task-specific adapter on top of the fixed l…

2022

Video-Text Representation Learning via Differentiable Weak Temporal Alignment

CVPR 2022poster

Learning generic joint representations for video and text by a supervised method requires a prohibitively substantial amount of manually annotated video datasets. As a practical alternative, a large-scale but uncurated and narrated video dataset, HowTo100M, has recently been introduced. But it is st…

Cited by 25PDFcodeScholar
2021

HOTR: End-to-End Human-Object Interaction Detection With Transformers

CVPR 2021poster

Human-Object Interaction (HOI) detection is a task of identifying "a set of interactions" in an image, which involves the i) localization of the subject (i.e., humans) and target (i.e., objects) of interaction, and ii) the classification of the interaction labels. Most existing methods have addresse…

Cited by 339PDFcodeScholar
2021

Image-to-Image Retrieval by Learning Similarity between Scene Graphs

AAAI 2021technical

As a scene graph compactly summarizes the high-level content of an image in a structured and symbolic manner, the similarity between scene graphs of two images reflects the relevance of their contents. Based on this idea, we propose a novel approach for image-to-image retrieval using scene graph sim…

2020

Hypergraph Attention Networks for Multimodal Learning

CVPR 2020poster

One of the fundamental problems that arise in multimodal learning tasks is the disparity of information levels between different modalities. To resolve this problem, we propose Hypergraph Attention Networks (HANs), which define a common semantic space among the modalities with symbolic graphs and ex…

Cited by 120PDFcodeScholar