← Search

Youngbeom Yoo

2 accepted papers

2025

Open-ended Hierarchical Streaming Video Understanding with Vision Language Models

ICCV 2025poster

We introduce Hierarchical Streaming Video Understanding, a task that combines online temporal action localization with free-form description generation. Given the scarcity of datasets with hierarchical and fine-grained temporal annotations, we demonstrate that LLMs can effectively group atomic actio…

Cited by 0SourcePDFScholar
2025

Representing 3D Shapes with 64 Latent Vectors for 3D Diffusion Models

ICCV 2025poster

Constructing a compressed latent space through a variational autoencoder (VAE) is the key for efficient 3D diffusion models. This paper introduces COD-VAE that encodes 3D shapes into a COmpact set of 1D latent vectors without sacrificing quality. COD-VAE introduces a two-stage autoencoder scheme to…