← Search

hongguang Zhu

4 accepted papers

2026

All-in-One Slider for Attribute Manipulation in Diffusion Models

CVPR 2026

Text-to-image (T2I) diffusion models have made significant strides in generating high-quality images. However, progressively manipulating certain attributes of generated images to meet the desired user expectations remains challenging, particularly for content with rich details, such as human faces.

Cited by 0SourcecodeScholar
2024

Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation

ECCV 2024oral

"Pre-trained vision-language models, e.g. CLIP, have been increasingly used to address the challenging Open-Vocabulary Segmentation (OVS) task, benefiting from their well-aligned vision-text embedding space. Typical solutions involve either freezing CLIP during training to unilaterally maintain its…

2023

CTP:Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation

ICCV 2023poster

Vision-Language Pretraining (VLP) has shown impressive results on diverse downstream tasks by offline training on large-scale datasets. Regarding the growing nature of real-world data, such an offline training paradigm on ever-expanding data is unsustainable, because models lack the continual learni…

Cited by 33PDFcodeScholar
2021

Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline

CVPR 2021poster

Depth maps obtained by commercial depth sensors are always in low-resolution, making it difficult to be used in various computer vision tasks. Thus, depth map super-resolution (SR) is a practical and valuable task, which upscales the depth map into high-resolution (HR) space. However, limited by the…

Cited by 100PDFScholar