← Search

Sihong Huang

2 accepted papers

2026

Compositional Transformation Reasoning for Composed Video Retrieval

CVPR 2026

Composed Video Retrieval aims to retrieve a target video given a reference video and a textual modification describing the desired change. The core challenge lies in modeling compositional multimodal transformations, i.e., how entities, actions, and scenes evolve across video and language modalities

Cited by 0SourcecodeScholar
2025

Sound Bridge: Associating Egocentric and Exocentric Videos via Audio Cues

CVPR 2025poster

Understanding human behavior and the environmental information in the egocentric video is very challenging due to the invisibility of some actions (e.g., laughing and sneezing) and the local nature of the first-person view. Leveraging the corresponding exocentric video to provide global context has…