← Search

Licheng Tang

3 accepted papers

2025

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

ACL 2025long

Multi-modal Large Language Models (MLLMs) have demonstrated remarkable reasoning capability while lack explicit mechanisms for visual grounding and segmentation, creating a gap between cognitive reasoning and visual perception. To bridge this gap, we introduce Reasoning Segmentation via Visual Promp…

Cited by 0SourcePDFScholar
2025

VEU-Bench: Towards Comprehensive Understanding of Video Editing

CVPR 2025highlight

Widely shared videos on the internet are often edited. Recently, although Video Large Language Models (Vid-LLMs) have made great progress in general video understanding tasks, their capabilities in video editing understanding (VEU) tasks remain unexplored. To address this gap, in this paper, we intr…

Cited by 0SourcePDFScholar
2022

Few-Shot Font Generation by Learning Fine-Grained Local Styles

CVPR 2022poster

Few-shot font generation (FFG), which aims to generate a new font with a few examples, is gaining increasing attention due to the significant reduction in labor cost. A typical FFG pipeline considers characters in a standard font library as content glyphs and transfers them to a new target font by e…

Cited by 80PDFcodeScholar