← Search

Junho Cho

3 accepted papers

2025

How Can Objects Help Video-Language Understanding?

ICCV 2025poster

Do we still need to represent objects explicitly in multimodal large language models (MLLMs)? To one extreme, pre-trained encoders convert images into visual tokens, with which objects and spatiotemporal relationships may be implicitly modeled. To the other extreme, image captions by themselves prov…

2025

SAGE: A Unified Framework for Generalizable Object State Recognition with State-Action Graph Embedding

NeurIPS 2025oral

Recognizing the physical states of objects and their transformations within videos is crucial for structured video understanding and enabling robust real-world applications, such as robotic manipulation. However, pretrained vision-language models often struggle to capture these nuanced dynamics and…

Cited by 0SourceScholar
2021

Unsupervised Hyperbolic Representation Learning via Message Passing Auto-Encoders

CVPR 2021poster

Most of the existing literature regarding hyperbolic embedding concentrate upon supervised learning, whereas the use of unsupervised hyperbolic embedding is less well explored. In this paper, we analyze how unsupervised tasks can benefit from learned representations in hyperbolic space. To explore h…

Cited by 39PDFcodeScholar