← Search

Jay Wu

3 accepted papers

2025

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

ACL 2025long

Multi-modal Large Language Models (MLLMs) have demonstrated remarkable reasoning capability while lack explicit mechanisms for visual grounding and segmentation, creating a gap between cognitive reasoning and visual perception. To bridge this gap, we introduce Reasoning Segmentation via Visual Promp…

Cited by 0SourcePDFScholar
2025

VEU-Bench: Towards Comprehensive Understanding of Video Editing

CVPR 2025highlight

Widely shared videos on the internet are often edited. Recently, although Video Large Language Models (Vid-LLMs) have made great progress in general video understanding tasks, their capabilities in video editing understanding (VEU) tasks remain unexplored. To address this gap, in this paper, we intr…

Cited by 0SourcePDFScholar
2025

Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark

AAAI 2025technical

The demand for producing short-form videos for sharing on social media platforms has experienced significant growth in recent times. Despite notable advancements in the fields of video summarization and highlight detection, which can create partially usable short films from raw videos, these approac…