← Search

Soo Won Seo

2 accepted papers

2026

Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection

CVPR 2026

Human-Object Interaction (HOI) detection aims to localize human-object pairs and classify their interactions from a single image, a task that demands strong visual understanding and nuanced contextual reasoning. Recent approaches have leveraged Vision-Language Models (VLMs) to introduce semantic pri

Cited by 0SourcecodeScholar
2025

JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts

AAAI 2025technical

Video Action Detection (VAD) entails localizing and categorizing action instances within videos, which inherently consist of diverse information sources such as audio, visual cues, and surrounding scene contexts. Leveraging this multi-modal information effectively for VAD poses a significant challen…