← Search

Insu Lee

3 accepted papers

2025

Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs

NeurIPS 2025spotlight

Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by head-mounted cameras serves as key input. While this view offers fine-grained cues about user attention and hand-object…

Cited by 0SourcecodeScholar
2024

Expand-and-Quantize: Unsupervised Semantic Segmentation Using High-Dimensional Space and Product Quantization

AAAI 2024technical

Unsupervised semantic segmentation (USS) aims to discover and recognize meaningful categories without any labels. For a successful USS, two key abilities are required: 1) information compression and 2) clustering capability. Previous methods have relied on feature dimension reduction for informatio…

Cited by 1SourcePDFScholar