← Search

Hongfa Wang

13 accepted papers

2026

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization

ICML 2026poster

Group Relative Policy Optimization has emerged as essential for aligning video diffusion models with human preferences, but faces a critical computational bottleneck: training a 14B parametered model typically demands hundreds of GPU days per experiment. Existing efficiency methods reduce costs thro…

Cited by 0SourceScholar
2026

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

AAAI 2026technical

Video captions play a crucial role in text-to-video generation tasks, as their quality directly influences the semantic coherence and visual fidelity of the generated videos. Although large vision-language models (VLMs) have demonstrated significant potential in caption generation, existing benchmar

Cited by 0SourcePDFScholar
2025

CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation

ICCV 2025poster

The newly proposed Generalized Referring Expression Segmentation (GRES) amplifies the formulation of classic RES by involving complex multiple/non-target scenarios. Recent approaches address GRES by directly extending the well-adopted RES frameworks with object-existence identification. However, the…

Cited by 0SourcePDFScholar
2025

Follow-Your-Click: Open-domain Regional Image Animation via Motion Prompts

AAAI 2025technical

Despite recent advances in image-to-video generation, better controllability and local animation are less explored. Most existing image-to-video methods are not locally aware and tend to move the entire scene. However, human artists may need to control the movement of different objects or regions. A…

Cited by 52SourcePDFScholar
2025

Infinite-Canvas: Higher-Resolution Video Outpainting with Extensive Content Generation

AAAI 2025technical

This paper explores higher-resolution video outpainting with extensive content generation. We point out common issues faced by existing methods when attempting to largely outpaint videos: the generation of low-quality content and limitations imposed by GPU memory. To address these challenges, we pro…

2025

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models

ICCV 2025poster

Boosted by Multi-modal Large Language Models (MLLMs), text-guided universal segmentation models for the image and video domains have made rapid progress recently. However, these methods are often developed separately for specific domains, overlooking the similarities in task settings and solutions a…

2025

Towards Multiple Character Image Animation Through Enhancing Implicit Decoupling

ICLR 2025poster

Controllable character image animation has a wide range of applications. Although existing studies have consistently improved performance, challenges persist in the field of character image animation, particularly concerning stability in complex backgrounds and tasks involving multiple characters. T…

Cited by 0SourcePDFScholar
2023

MAP: Multimodal Uncertainty-Aware Vision-Language Pre-Training Model

CVPR 2023poster

Multimodal semantic understanding often has to deal with uncertainty, which means the obtained messages tend to refer to multiple targets. Such uncertainty is problematic for our interpretation, including inter- and intra-modal uncertainty. Little effort has studied the modeling of this uncertainty,…

2023

Seeing What You Miss: Vision-Language Pre-Training With Semantic Completion Learning

CVPR 2023poster

Cross-modal alignment is essential for vision-language pre-training (VLP) models to learn the correct corresponding information across different modalities. For this purpose, inspired by the success of masked language modeling (MLM) tasks in the NLP pre-training area, numerous masked modeling tasks…

2022

Deep Unsupervised Hashing with Latent Semantic Components

AAAI 2022technical

Deep unsupervised hashing has been appreciated in the regime of image retrieval. However, most prior arts failed to detect the semantic components and their relationships behind the images, which makes them lack discriminative power. To make up the defect, we propose a novel Deep Semantic Component…

Cited by 26SourcePDFScholar
2021

Adaptive Boundary Proposal Network for Arbitrary Shape Text Detection

ICCV 2021poster

Arbitrary shape text detection is a challenging task due to the high complexity and variety of scene texts. In this work, we propose a novel adaptive boundary proposal network for arbitrary shape text detection, which can learn to directly produce accurate boundary for arbitrary shape text without a…

Cited by 121PDFcodeScholar
2020

Deep Relational Reasoning Graph Network for Arbitrary Shape Text Detection

CVPR 2020oral

Arbitrary shape text detection is a challenging task due to the high variety and complexity of scenes texts. In this paper, we propose a novel unified relational reasoning graph network for arbitrary shape text detection. In our method, an innovative local graph bridges a text proposal model via Con…

Cited by 281PDFcodeScholar