← Search

Dian Zheng

12 accepted papers

2026

EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding

AAAI 2026technical

In this work, we present EchoGen, a unified framework for layout-to-image generation and image grounding, capable of generating images with both accurate layout and high fidelity to the text description.(e.g., spatial relationship), and grounding the image robustly at the same time. We believe that

Cited by 0SourcePDFScholar
2026

OneThinker: All-in-one Reasoning Model for Image and Video

CVPR 2026

Reinforcement learning (RL) has recently achieved remarkable success in eliciting visual reasoning within Multimodal Large Language Models (MLLMs). However, existing approaches typically train separate models for different tasks and treat image and video reasoning as disjoint domains. This results i

Cited by 0SourcecodeScholar
2025

Decoupled Distillation to Erase: A General Unlearning Method for Any Class-centric Tasks

CVPR 2025highlight

In this work, we present DEcoupLEd Distillation To Erase (DELETE), a general and strong unlearning method for any class-centric tasks. To derive this, we first propose a theoretical framework to analyze the general form of unlearning loss and decompose it into forgetting and retention terms. Through…

Cited by 2SourcePDFScholar
2025

Panorama Generation From NFoV Image Done Right

CVPR 2025highlight

Generating 360-degree panoramas from narrow field of view (NFoV) image is a promising computer vision task for Virtual Reality (VR) applications. Existing methods mostly assess the generated panoramas with InceptionNet or CLIP based metrics, which tend to perceive the image quality and is not suitab…

2025

ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models

NeurIPS 2025poster

Recent Vision-Language Models (VLMs) have shown strong performance in general-purpose visual understanding and reasoning, but their ability to comprehend the visual grammar of movie shots remains underexplored and insufficiently evaluated. To bridge this gap, we present \textbf{ShotBench}, a dedicat…

Cited by 0SourceScholar
2025

SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular Input

CVPR 2025poster

Stereo video synthesis from monocular input is challenging in spatial computing and virtual reality due to the lack of high-quality stereo video pairs for training and the difficulty of maintaining spatio-temporal consistency between frames. Existing methods primarily address these issues by directl…

2024

An Economic Framework for 6-DoF Grasp Detection

ECCV 2024poster

"Robotic grasping in clutters is a fundamental task in robotic manipulation. In this work, we propose an economic framework for 6-DoF grasp detection, aiming to economize the resource cost in training and meanwhile maintain effective grasp performance. To begin with, we discover that the dense super…

2024

Selective Hourglass Mapping for Universal Image Restoration Based on Diffusion Model

CVPR 2024poster

Universal image restoration is a practical and potential computer vision task for real-world applications. The main challenge of this task is handling the different degradation distributions at once. Existing methods mainly utilize task-specific conditions (e.g. prompt) to guide the model to learn d…

2023

Estimator Meets Equilibrium Perspective: A Rectified Straight Through Estimator for Binary Neural Networks Training

ICCV 2023poster

Binarization of neural networks is a dominant paradigm in neural networks compression. The pioneering work BinaryConnect uses Straight Through Estimator (STE) to mimic the gradients of the sign function, but it also causes the crucial inconsistency problem. Most of the previous methods design differ…

Cited by 19PDFcodeScholar
2023

Generating Anomalies for Video Anomaly Detection With Prompt-Based Feature Mapping

CVPR 2023poster

Anomaly detection in surveillance videos is a challenging computer vision task where only normal videos are available during training. Recent work released the first virtual anomaly detection dataset to assist real-world detection. However, an anomaly gap exists because the anomalies are bounded in…

Cited by 42SourcePDFScholar
2022

Underwater Stereo Matching Via Unsupervised Appearance And Feature Adaptation Networks

ICASSP 2022accepted

Stereo matching has been widely used to estimate depth maps in terrestrial environments. However, it is difficult to achieve appealing performance in underwater environments, since adequate underwater stereo data with groundtruth depth information is not easily available for training an underwater d…

Cited by 0SourceScholar