← Search

Minyoung Noh

1 accepted papers

2025

Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs

NeurIPS 2025spotlight

Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by head-mounted cameras serves as key input. While this view offers fine-grained cues about user attention and hand-object…

Cited by 0SourcecodeScholar