← Search

Zhouxia Wang

10 accepted papers

2026

Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation

CVPR 2026

Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlusions and during cross-task transfer. To address these challenges, we propose Task-aware Virtual View Exploration (TVVE)

Cited by 0SourcecodeScholar
2026

Precise Object and Effect Removal with Adaptive Target-Aware Attention

CVPR 2026

Object removal requires eliminating not only the target object but also its associated visual effects such as shadows and reflections. However, diffusion-based inpainting and removal methods often introduce artifacts, hallucinate contents, alter background, and struggle to remove object effects accu

Cited by 0SourcecodeScholar
2025

Denoising as Adaptation: Noise-Space Domain Adaptation for Image Restoration

ICLR 2025poster

Although learning-based image restoration methods have made significant progress, they still struggle with limited generalization to real-world scenarios due to the substantial domain gap caused by training on synthetic data. Existing methods address this issue by improving data synthesis pipelines,…

2025

Image Conductor: Precision Control for Interactive Video Synthesis

AAAI 2025technical

Filmmaking and animation production often require sophisticated techniques for coordinating camera transitions and object movements, typically involving labor-intensive real-world capturing. Despite advancements in generative AI for video creation, achieving precise control over motion for interacti…

2024

Diffusion-based Blind Text Image Super-Resolution

CVPR 2024poster

Recovering degraded low-resolution text images is challenging especially for Chinese text images with complex strokes and severe degradation in real-world scenarios. Ensuring both text fidelity and style realness is crucial for high-quality text image super-resolution. Recently diffusion models have…

2022

RestoreFormer: High-Quality Blind Face Restoration From Undegraded Key-Value Pairs

CVPR 2022poster

Blind face restoration is to recover a high-quality face image from unknown degradations. As face image contains abundant contextual information, we propose a method, RestoreFormer, which explores fully-spatial attentions to model contextual information and surpasses existing works that use local co…

Cited by 123PDFcodeScholar
2020

Learning a Reinforced Agent for Flexible Exposure Bracketing Selection

CVPR 2020poster

Automatically selecting exposure bracketing (images exposed differently) is important to obtain a high dynamic range image by using multi-exposure fusion. Unlike previous methods that have many restrictions such as requiring camera response function, sensor noise model, and a stream of preview image…

Cited by 24PDFcodeScholar
2017

Multi-Label Image Recognition by Recurrently Discovering Attentional Regions

ICCV 2017poster

This paper proposes a novel deep architecture to address multi-label image recognition, a fundamental and practical task towards general visual understanding. Current solutions for this task usually rely on an extra step of extracting hypothesis regions (i.e., region proposals), resulting in redunda…

Cited by 394PDFScholar