← Search

Zheng Gu

6 accepted papers

2026

Cycle-Consistent Tuning for Layered Image Decomposition

CVPR 2026

Disentangling visual layers in real-world images is a persistent challenge in vision and graphics, as such layers often involve non-linear and globally coupled interactions, including shading, reflection, and perspective distortion. In this work, we present an in-context image decomposition framewor

Cited by 0SourcecodeScholar
2026

DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video

CVPR 2026

Reliable 4D object detection, which refers to 3D object detection in streaming video, is crucial for perceiving and understanding the real world. Existing open-set 4D object detection methods typically make predictions on a frame-by-frame basis without modeling temporal consistency, or rely on compl

Cited by 0SourcecodeScholar
2025

CoT-VTM: Visual-to-Music Generation with Chain-of-Thought Reasoning

ACL 2025finding

The application of visual-to-music generation (VTM) is rapidly growing. However, current VTM methods struggle with capturing the relationship between visuals and music in open-domain settings, mainly due to two challenges: the lack of large-scale, high-quality visual-music paired datasets and the ab…

2021

LoFGAN: Fusing Local Representations for Few-Shot Image Generation

ICCV 2021poster

Given only a few available images for a novel unseen category, few-shot image generation aims to generate more data for this category. Previous works attempt to globally fuse these images by using adjustable weighted coefficients. However, there is a serious semantic misalignment between different i…

Cited by 79PDFcodeScholar
2020

Learning Task-aware Local Representations for Few-shot Learning

IJCAI 2020poster

Few-shot learning for visual recognition aims to adapt to novel unseen classes with only a few images. Recent work, especially the work based on low-level information, has achieved great progress. In these work, local representations (LRs) are typically employed, because LRs are more consistent amon…

Cited by 0SourcePDFScholar
2020

Unsupervised Domain Attention Adaptation Network for Caricature Attribute Recognition

ECCV 2020poster

Caricature attributes provide distinctive facial features to help research in Psychology and Neuroscience. However, unlike the facial photo attribute datasets that have a quantity of annotated images, the annotations of caricature attributes are rare. To facility the research in attribute learning o…