← Search

Jihoon Chung

4 accepted papers

2025

Unifying Specialized Visual Encoders for Video Language Models

ICML 2025poster

Recent advances in vision backbones have yielded powerful and diverse visual and video encoders. Yet, current Video Large Language Models encode visual inputs using an encoder from a single backbone family, limiting the amount and type of visual information they can process. We propose MERV, a Multi…

2022

Enabling Detailed Action Recognition Evaluation Through Video Dataset Augmentation

NeurIPS 2022accept

It is well-known in the video understanding community that human action recognition models suffer from background bias, i.e., over-relying on scene cues in making their predictions. However, it is difficult to quantify this effect using existing evaluation frameworks. We introduce the Human-centric…

Cited by 14SourcePDFScholar
2021

HAA500: Human-Centric Atomic Action Dataset With Curated Videos

ICCV 2021poster

We contribute HAA500, a manually annotated human-centric atomic action dataset for action recognition on 500 classes with over 591K labeled frames. To minimize ambiguities in action classification, HAA500 consists of highly diversified classes of fine-grained atomic actions, where only consistent ac…

Cited by 60PDFScholar
2020

CascadePSP: Toward Class-Agnostic and Very High-Resolution Segmentation via Global and Local Refinement

CVPR 2020poster

State-of-the-art semantic segmentation methods were almost exclusively trained on images within a fixed resolution range. These segmentations are inaccurate for very high-resolution images since using bicubic upsampling of low-resolution segmentation does not adequately capture high-resolution detai…

Cited by 289PDFcodeScholar