← Search

Jaewook Yoo

3 accepted papers

2025

How Can Objects Help Video-Language Understanding?

ICCV 2025poster

Do we still need to represent objects explicitly in multimodal large language models (MLLMs)? To one extreme, pre-trained encoders convert images into visual tokens, with which objects and spatiotemporal relationships may be implicitly modeled. To the other extreme, image captions by themselves prov…

2025

SAGE: A Unified Framework for Generalizable Object State Recognition with State-Action Graph Embedding

NeurIPS 2025oral

Recognizing the physical states of objects and their transformations within videos is crucial for structured video understanding and enabling robust real-world applications, such as robotic manipulation. However, pretrained vision-language models often struggle to capture these nuanced dynamics and…

Cited by 0SourceScholar
2022

SeeThroughNet: Resurrection of Auxiliary Loss by Preserving Class Probability Information

CVPR 2022poster

Auxiliary loss is additional loss besides the main branch loss to help optimize the learning process of neural networks. In order to calculate the auxiliary loss between the feature maps of intermediate layers and the ground truth in the field of semantic segmentation, the size of each feature map m…

Cited by 7PDFScholar