← Search

Dingwen Zhang

37 accepted papers

2026

Few-Shot Hybrid Incremental Learning:Continually Learning under Data Scarcity and Task Uncertainty

CVPR 2026

The increasing complexity of real-world deployment requires intelligent agents to effectively adapt to non-stationary data streams with stochastic increments under data scarcity. We formally define this challenge as the Few-Shot Hybrid Incremental Learning (FSHIL) paradigm, which reveals a critical

Cited by 0SourceScholar
2026

GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic Characteristics

CVPR 2026

This paper presents GeoAgent, a model capable of reasoning closely with humans and deriving fine-grained address conclusions. Previous RL-based methods have achieved breakthroughs in performance and interpretability but still remain concerns because of their reliance on AI-generated chain-of-thought

Cited by 0SourceScholar
2026

Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence

ICML 2026oral

The pursuit of spatial intelligence fundamentally relies on access to large-scale, fine-grained 3D data. However, existing approaches predominantly construct spatial understanding benchmarks by generating question–answer (QA) pairs from a limited number of manually annotated datasets, rather than sy…

Cited by 0SourceScholar
2026

IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction

ICLR 2026poster

Humans naturally perceive the geometric structure and semantic content of a 3D world as intertwined dimensions, enabling coherent and accurate understanding of complex scenes. However, most prior approaches prioritize training large geometry models for low-level 3D reconstruction and treat high-leve…

Cited by 0SourcecodeScholar
2026

Proxy-GS: Unified Occlusion Priors for Training and Inference in Structured 3D Gaussian Splatting

CVPR 2026

3D Gaussian Splatting (3DGS) has emerged as an efficient approach for photorealistic rendering. Recent MLP-based variants further improve visual fidelity but introduce substantial decoding overhead during rendering. To reduce the computational cost, several pruning strategies and level-of-detail (LO

Cited by 0SourcecodeScholar
2025

CityGS-X: A Scalable Architecture for Efficient and Geometrically Accurate Large-Scale Scene Reconstruction

ICCV 2025poster

Despite its significant achievements in large-scale scene reconstruction, 3D Gaussian Splatting still faces substantial challenges, including slow processing, high computational costs, and limited geometric accuracy. These core issues arise from its inherently unstructured design and the absence of…

Cited by 0SourcePDFScholar
2025

DGTR: Distributed Gaussian Turbo-Reconstruction for Sparse-View Vast Scenes

ICRA 2025

Novel-view synthesis approaches play a critical role in vast scene reconstruction. However, these methods rely heavily on dense image inputs and prolonged training times, making them unsuitable where computational resources are limited. Additionally, few-shot methods often struggle with poor reconst

Cited by 5SourcecodeScholar
2025

Navigating Semantic Drift in Task-Agnostic Class-Incremental Learning

ICML 2025oral

Class-incremental learning (CIL) seeks to enable a model to sequentially learn new classes while retaining knowledge of previously learned ones. Balancing flexibility and stability remains a significant challenge, particularly when the task ID is unknown. To address this, our study reveals that the…

2025

Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction

CVPR 2025poster

Video virtual try-on aims to seamlessly dress a subject in a video with a specific garment. The primary challenge involves preserving the visual authenticity of the garment while dynamically adapting to the pose and physique of the subject. While existing methods have predominantly focused on image-…

Cited by 0SourcePDFScholar
2025

STRIDER: Navigation via Instruction-Aligned Structural Decision Space Optimization

NeurIPS 2025poster

The Zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) task requires agents to navigate previously unseen 3D environments using natural language instructions, without any scene-specific training. A critical challenge in this setting lies in ensuring agents’ actions align wi…

Cited by 0SourceScholar
2025

VDG: Vision-Only Dynamic Gaussian for Driving Simulation

RA-L 2025

Recent advances in dynamic Gaussian splatting have significantly improved scene reconstruction and novel-view synthesis. However, existing methods often rely on pre-computed camera poses and Gaussian initialization using Structure from Motion (SfM) or other costly sensors, limiting their scalability

Cited by 23SourceScholar
2024

CONDA: Condensed Deep Association Learning for Co-Salient Object Detection.

ECCV 2024poster

"Inter-image association modeling is crucial for co-salient object detection. Despite satisfactory performance, previous methods still have limitations on sufficient inter-image association modeling. Because most of them focus on image feature optimization under the guidance of heuristically calcula…

2024

GP-NeRF: Generalized Perception NeRF for Context-Aware 3D Scene Understanding

CVPR 2024highlight

Applying Neural Radiance Fields (NeRF) to downstream perception tasks for scene understanding and representation is becoming increasingly popular. Most existing methods treat semantic prediction as an additional rendering task i.e. the "label rendering" task to build semantic NeRFs. However by rende…

Cited by 25SourcePDFScholar
2024

Gradient and Brightness Guided Low-Light Enhancement with Attention-Based Self-Paced Learning

ICASSP 2024accepted

Low-light image enhancement aims to reconstruct images with insufficient illumination into visually appealing representations with natural brightness. While most existing methods tend to focus on enhancing illumination, they often overlook the restoration of finer details in the enhanced image. More…

Cited by 0SourceScholar
2024

Revisiting the Power of Prompt for Visual Tuning

ICML 2024spotlight

Visual prompt tuning (VPT) is a promising solution incorporating learnable prompt tokens to customize pre-trained models for downstream tasks. However, VPT and its variants often encounter challenges like prompt initialization, prompt length, and subpar performance in self-supervised pretraining, hi…

2024

SpFormer: Spatio-Temporal Modeling for Scanpaths with Transformer

AAAI 2024technical

Saccadic scanpath, a data representation of human visual behavior, has received broad interest in multiple domains. Scanpath is a complex eye-tracking data modality that includes the sequences of fixation positions and fixation duration, coupled with image information. However, previous methods usua…

2024

Task-aware Orthogonal Sparse Network for Exploring Shared Knowledge in Continual Learning

ICML 2024poster

Continual learning (CL) aims to learn from sequentially arriving tasks without catastrophic forgetting (CF). By partitioning the network into two parts based on the Lottery Ticket Hypothesis---one for holding the knowledge of the old tasks while the other for learning the knowledge of the new task--…

Cited by 7SourcePDFScholar
2023

Boosting Low-Data Instance Segmentation by Unsupervised Pre-Training With Saliency Prompt

CVPR 2023poster

Recently, inspired by DETR variants, query-based end-to-end instance segmentation (QEIS) methods have outperformed CNN-based models on large-scale datasets. Yet they would lose efficacy when only a small amount of training data is available since it's hard for the crucial queries/kernels to learn lo…

2022

Incremental Cross-View Mutual Distillation for Self-Supervised Medical CT Synthesis

CVPR 2022poster

Due to the constraints of the imaging device and high cost in operation time, computer tomography (CT) scans are usually acquired with low within-slice resolution. Improving the inter-slice resolution is beneficial to the disease diagnosis for both human experts and computer-aided systems. To this e…

Cited by 25PDFScholar
2022

Robust Region Feature Synthesizer for Zero-Shot Object Detection

CVPR 2022poster

Zero-shot object detection aims at incorporating class semantic vectors to realize the detection of (both seen and) unseen classes given an unconstrained test image. In this study, we reveal the core challenges in this research area: how to synthesize robust region features (for unseen objects) that…

Cited by 54PDFcodeScholar
2022

Robust Single Image Dehazing Based on Consistent and Contrast-Assisted Reconstruction

IJCAI 2022poster

Single image dehazing as a fundamental low-level vision task, is essential for the development of robust intelligent surveillance system. In this paper, we make an early effort to consider dehazing robustness under variational haze density, which is a realistic while under-studied problem in the res…

Cited by 7SourcePDFScholar
2021

ABMDRNet: Adaptive-Weighted Bi-Directional Modality Difference Reduction Network for RGB-T Semantic Segmentation

CVPR 2021poster

Semantic segmentation models gain robustness against poor lighting conditions by virtue of complementary information from visible (RGB) and thermal images. Despite its importance, most existing RGB-T semantic segmentation models perform primitive fusion strategies, such as concatenation, element-wis…

Cited by 178PDFScholar
2021

Light Field Saliency Detection With Dual Local Graph Learning and Reciprocative Guidance

ICCV 2021poster

The application of light field data in salient object detection is becoming increasingly popular in recent years. The difficulty lies in how to effectively fuse the features within the focal stack and how to cooperate them with the feature of the all-focus image. Previous methods usually fuse focal…

Cited by 48PDFcodeScholar
2021

Strengthen Learning Tolerance for Weakly Supervised Object Localization

CVPR 2021poster

Weakly supervised object localization (WSOL) aims at learning to localize objects of interest by only using the image-level labels as the supervision. While numerous efforts have been made in this field, recent approaches still suffer from two challenges: one is the part domination issue while the o…

Cited by 77PDFcodeScholar
2020

Few-Cost Salient Object Detection with Adversarial-Paced Learning

NeurIPS 2020poster

Detecting and segmenting salient objects from given image scenes has received great attention in recent years. A fundamental challenge in training the existing deep saliency detection models is the requirement of large amounts of annotated data. While gathering large quantities of training data beco…

2020

Taking a Deeper Look at Co-Salient Object Detection

CVPR 2020poster

Co-salient object detection (CoSOD) is a newly emerging and rapidly growing branch of salient object detection (SOD), which aims to detect the co-occurring salient objects in multiple images. However, existing CoSOD datasets often have a serious data bias, which assumes that each group of images con…

Cited by 100PDFScholar
2018

PoseFlow: A Deep Motion Representation for Understanding Human Behaviors in Videos

CVPR 2018poster

Motion of the human body is the critical cue for understanding and characterizing human behavior in videos. Most existing approaches explore the motion cue using optical flows. However, optical flow usually contains motion on both the interested human bodies and the undesired background. This "noisy…

Cited by 43SourcePDFScholar
2018

Reinforcement Cutting-Agent Learning for Video Object Segmentation

CVPR 2018poster

Video object segmentation is a fundamental yet challenging task in computer vision community. In this paper, we formulate this problem as a Markov Decision Process, where agents are learned to segment object regions under a deep reinforcement learning framework. Essentially, learning agents for segm…

Cited by 105SourcePDFScholar
2017

SPFTN: A Self-Paced Fine-Tuning Network for Segmenting Objects in Weakly Labelled Videos

CVPR 2017poster

Object segmentation in weakly labelled videos is an interesting yet challenging task, which aims at learning to perform category-specific video object segmentation by only using video-level tags. Existing works in this research area might still have some limitations, e.g., lack of effective DNN-base…

Cited by 61PDFScholar
2015

A Self-Paced Multiple-Instance Learning Framework for Co-Saliency Detection

ICCV 2015poster

As an interesting and emerging topic, co-saliency detection aims at simultaneously extracting common salient objects in a group of images. Traditional co-saliency detection approaches rely heavily on human knowledge for designing hand-crafted metrics to explore the intrinsic patterns underlying co-s…

Cited by 155PDFScholar
2015

Predicting Eye Fixations Using Convolutional Neural Networks

CVPR 2015poster

It is believed that eye movements in free-viewing of natural scenes are directed by both bottom-up visual saliency and top-down visual factors. In this paper, we propose a novel computational framework to simultaneously learn these two types of visual features from raw image data using a multiresolu…