← Search

Dahun Kim

17 accepted papers

2025

VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models

CVPR 2025poster

We introduce VideoComp, a benchmark and learning framework for advancing video-text compositionality understanding, aimed at improving vision-language models (VLMs) in fine-grained temporal alignment. Unlike existing benchmarks focused on static image-text compositionality or isolated single-event v…

2024

Mirasol3B: A Multimodal Autoregressive Model for Time-Aligned and Contextual Modalities

CVPR 2024poster

One of the main challenges of multimodal learning is the need to combine heterogeneous modalities (e.g. video audio text). For example video and audio are obtained at much higher rates than text and are roughly aligned in time. They are often not synchronized with text which comes as a global contex…

Cited by 23SourcePDFScholar
2024

Uni-DVPS: Unified Model for Depth-Aware Video Panoptic Segmentation

RA-L 2024

We present Uni-DVPS, a unified model for Depth-aware Video Panoptic Segmentation (DVPS) that jointly tackles distinct vision tasks, i.e., video panoptic segmentation, monocular depth estimation, and object tracking. In contrast to the prior works that adopt diverged decoder networks tailored for eac

Cited by 6SourceScholar
2023

Neural Image-based Avatars: Generalizable Radiance Fields for Human Avatar Modeling

ICLR 2023poster

We present a method that enables synthesizing novel views and novel poses of arbitrary human performers from sparse multi-view images. A key ingredient of our method is a hybrid appearance blending module that combines the advantages of the implicit body NeRF representation and image-based rendering…

Cited by 19SourcePDFScholar
2023

Region-Aware Pretraining for Open-Vocabulary Object Detection With Vision Transformers

CVPR 2023highlight

We present Region-aware Open-vocabulary Vision Transformers (RO-ViT) -- a contrastive image-text pretraining recipe to bridge the gap between image-level pretraining and open-vocabulary object detection. At the pretraining phase, we propose to randomly crop and resize regions of positional embedding…

Cited by 84SourcePDFScholar
2022

CMT-DeepLab: Clustering Mask Transformers for Panoptic Segmentation

CVPR 2022oral

We propose Clustering Mask Transformer (CMT-DeepLab), a transformer-based framework for panoptic segmentation designed around clustering. It rethinks the existing transformer architectures used in segmentation and detection; CMT-DeepLab considers the object queries as cluster centers, which fill the…

Cited by 110PDFScholar
2022

Learning Open-World Object Proposals Without Learning to Classify

RA-L 2022

Object proposals have become an integral pre-processing step of many vision pipelines including object detection, weakly supervised detection, object discovery, tracking, etc. Compared to the learning-free methods, learning-based proposals have become popular recently due to the growing interest in

Cited by 158SourcecodeScholar
2021

Neural Human Performer: Learning Generalizable Radiance Fields for Human Performance Rendering

NeurIPS 2021spotlight

In this paper, we aim at synthesizing a free-viewpoint video of an arbitrary human performance using sparse multi-view cameras. Recently, several works have addressed this problem by learning person-specific neural radiance fields (NeRF) to capture the appearance of a particular human. In parallel,…

2020

Rotationally-Temporally Consistent Novel View Synthesis of Human Performance Video

ECCV 2020poster

Novel view video synthesis aims to synthesize novel viewpoints videos given input captures of a human performance taken from multiple reference viewpoints and over consecutive time steps. Despite great advances in model-free novel view synthesis, existing methods present three limitations when appli…

Cited by 18SourcePDFScholar