← Search

Bihan Wen

53 accepted papers

2026

Conformal Prediction for Multi-Source Detection on a Network

AAAI 2026technical

Detecting the origin of information or infection spread in networks is a fundamental challenge with applications in misinformation tracking, epidemiology, and beyond. We study the multi-source detection problem: given snapshot observations of node infection status on a graph, estimate the set of sou

Cited by 0SourcePDFScholar
2026

Phy-CoSF: Physics-Guided Continuous Spectral Fields Reconstruction and Spectral Super-Resolution for Snapshot Compressive Imaging

ICML 2026poster

Recent advances have demonstrated that coded aperture snapshot spectral imaging (CASSI) systems show great potential for capturing 3D hyperspectral images (HSIs) from a single 2D measurement. Despite the inherent spectral continuity of scenes captured by CASSI, most existing reconstruction methods a…

Cited by 0SourceScholar
2026

SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision

AAAI 2026technical

Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previous mitigation methods effectively reduce hallucinations in photographic images, t

Cited by 0SourcePDFScholar
2025

Discovering Hidden Visual Concepts Beyond Linguistic Input in Infant Learning

CVPR 2025poster

Infants develop complex visual understanding rapidly, even preceding of the acquisition of linguistic skills. As computer vision seeks to replicate the human vision system, understanding infant visual development may offer valuable insights. In this paper, we present an interdisciplinary study explo…

2025

Distribution Alignment Informed Thresholding for Semi-Supervised Curvilinear Structure Segmentation

ICASSP 2025accepted

Curvilinear structure segmentation using deep neural networks is often limited by the high cost of annotation. Semi-supervised learning (SSL) helps mitigate this dependency on extensive annotated data. State-of-the-art SSL approaches generate pseudo-labels for unlabeled data, which are then used for…

Cited by 0SourceScholar
2025

GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding

ICCV 2025poster

Sequential grounding in 3D point clouds (SG3D) refers to locating sequences of objects by following text instructions for a daily activity with detailed steps. Current 3D visual grounding (3DVG) methods treat text instructions with multiple steps as a whole, without extracting useful temporal inform…

Cited by 0SourcePDFScholar
2025

KaRF: Weakly-Supervised Kolmogorov-Arnold Networks-based Radiance Fields for Local Color Editing

NeurIPS 2025poster

Recent advancements have suggested that neural radiance fields (NeRFs) show great potential in color editing within the 3D domain. However, most existing NeRF-based editing methods continue to face significant challenges in local region editing, which usually lead to imprecise local object boundarie…

Cited by 0SourcecodeScholar
2025

M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision

ICCV 2025poster

RGB-Thermal (RGBT) multispectral vision is essential for robust perception in complex environments. Most RGBT tasks follow a case-by-case research paradigm, relying on manually customized models to learn task-oriented representations. Nevertheless, this paradigm is inherently constrained by artifici…

Cited by 0SourcePDFScholar
2025

Reconciling Stochastic and Deterministic Strategies for Zero-shot Image Restoration using Diffusion Model in Dual

CVPR 2025poster

Plug-and-play (PnP) methods offer an iterative strategy for solving image restoration (IR) problems in a zero-shot manner, using a learned discriminative denoiser as the implicit prior. More recently, a sampling-based variant of this approach, which utilizes a pre-trained generative diffusion model,…

2025

SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance Segmentation

ICCV 2025poster

Most existing remote sensing instance segmentation approaches are designed for close-vocabulary prediction, limiting their ability to recognize novel categories or generalize across datasets. This restricts their applicability in diverse Earth observation scenarios. To address this, we introduce ope…

2025

SoftShadow: Leveraging Soft Masks for Penumbra-Aware Shadow Removal

CVPR 2025poster

Recent advancements in deep learning have yielded promising results for the image shadow removal task. However, most existing methods rely on binary pre-generated shadow masks. The binary nature of such masks could potentially lead to artifacts near the boundary between shadow and non-shadow areas.…

Cited by 0SourcePDFScholar
2025

Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes Modeling

NeurIPS 2025poster

High-fidelity 3D object synthesis remains significantly more challenging than 2D image generation due to the unstructured nature of mesh data and the cubic complexity of dense volumetric grids. Existing two-stage pipelines—compressing meshes with a VAE (using either 2D or 3D supervision), followed b…

Cited by 0SourceScholar
2025

Training-Free Text-Guided Image Editing with Visual Autoregressive Model

ICCV 2025poster

Text-guided image editing is an essential task, enabling users to modify images through natural language descriptions. Recent advances in diffusion models and rectified flows have significantly improved editing quality, primarily relying on inversion techniques to extract structured noise from input…

2025

ZoRI: Towards Discriminative Zero-Shot Remote Sensing Instance Segmentation

AAAI 2025technical

Instance segmentation algorithms in remote sensing are typically based on conventional methods, limiting their application to seen scenarios and closed-set predictions. In this work, we propose a novel task called zero-shot remote sensing instance segmentation, aimed at identifying aerial objects th…

2024

Benchmarking Adversarial Robustness of Image Shadow Removal with Shadow-Adaptive Attacks

ICASSP 2024accepted

Shadow removal is a task aimed at erasing regional shadows present in images and reinstating visually pleasing natural scenes with consistent illumination. While recent deep learning techniques have demonstrated impressive performance in image shadow removal, their robustness against adversarial att…

Cited by 0SourceScholar
2024

CaKDP: Category-aware Knowledge Distillation and Pruning Framework for Lightweight 3D Object Detection

CVPR 2024poster

Knowledge distillation (KD) possesses immense potential to accelerate the deep neural networks (DNNs) for LiDAR-based 3D detection. However in most of prevailing approaches the suboptimal teacher models and insufficient student architecture investigations limit the performance gains. To address thes…

2024

Compress Clean Signal from Noisy Raw Image: A Self-Supervised Approach

ICML 2024poster

Raw images offer unique advantages in many low-level visual tasks due to their unprocessed nature. However, this unprocessed state accentuates noise, making raw images challenging to compress effectively. Current compression methods often overlook the ubiquitous noise in raw space, leading to increa…

Cited by 0SourcePDFScholar
2024

ContextGS : Compact 3D Gaussian Splatting with Anchor Level Context Model

NeurIPS 2024poster

Recently, 3D Gaussian Splatting (3DGS) has become a promising framework for novel view synthesis, offering fast rendering speeds and high fidelity. However, the large number of Gaussians and their associated attributes require effective compression techniques. Existing methods primarily compress ne…

2024

Joint RGB-Spectral Decomposition Model Guided Image Enhancement in Mobile Photography

ECCV 2024poster

"The integration of miniaturized spectrometers into mobile devices offers new avenues for image quality enhancement and facilitates novel downstream tasks. However, the broader application of spectral sensors in mobile photography is hindered by the inherent complexity of spectral images and the con…

2024

Progressive Divide-and-Conquer via Subsampling Decomposition for Accelerated MRI

CVPR 2024highlight

Deep unfolding networks (DUN) have emerged as a popular iterative framework for accelerated magnetic resonance imaging (MRI) reconstruction. However conventional DUN aims to reconstruct all the missing information within the entire space in each iteration. Thus it could be challenging when dealing w…

2024

SinSR: Diffusion-Based Image Super-Resolution in a Single Step

CVPR 2024poster

While super-resolution (SR) methods based on diffusion models exhibit promising results their practical application is hindered by the substantial number of required inference steps. Recent methods utilize the degraded images in the initial state thereby shortening the Markov chain. Nevertheless the…

2024

Temporal As a Plugin: Unsupervised Video Denoising with Pre-Trained Image Denoisers

ECCV 2024poster

"Recent advancements in deep learning have shown impressive results in image and video denoising, leveraging extensive pairs of noisy and noise-free data for supervision. However, the challenge of acquiring paired videos for dynamic scenes hampers the practical deployment of deep video denoising tec…

2024

Video-Text Prompting for Weakly Supervised Spatio-Temporal Video Grounding

EMNLP 2024main

Weakly-supervised Spatio-Temporal Video Grounding(STVG) aims to localize target object tube given a text query, without densely annotated training data. Existing methods extract each candidate tube feature independently by cropping objects from video frame feature, discarding all contextual informat…

Cited by 0SourcePDFScholar
2023

Boundary-Aware Divide and Conquer: A Diffusion-Based Solution for Unsupervised Shadow Removal

ICCV 2023poster

Recent deep learning methods have achieved superior results in shadow removal. However, most of these supervised methods rely on training over a huge amount of shadow and shadow-free image pairs, which require laborious annotations and may end up with poor model generalization. Shadows, in fact, onl…

Cited by 19PDFScholar
2023

ExposureDiffusion: Learning to Expose for Low-light Image Enhancement

ICCV 2023poster

Previous raw image-based low-light image enhancement methods predominantly relied on feed-forward neural networks to learn deterministic mappings from low-light to normally-exposed images. However, they failed to capture critical distribution information, leading to visually undesirable results. Thi…

Cited by 65PDFcodeScholar
2023

Hyperspectral Image Denoising Via Nonlocal Rank Residual Modeling

ICASSP 2023accepted

Nonlocal low-rank (LR) tensor modeling has shown great potential in hyperspectral image (HSI) denoising, which first uses the nonlocal self-similarity (NSS) prior to search for many similar full-band patches to form three-dimensional nonlocal full-band groups (tensors), and then usually enforces an…

Cited by 0SourceScholar
2023

Raw Image Reconstruction With Learned Compact Metadata

CVPR 2023poster

While raw images exhibit advantages over sRGB images (e.g. linearity and fine-grained quantization level), they are not widely used by common users due to the large storage requirements. Very recent works propose to compress raw images by designing the sampling masks in the raw image pixel space, le…

2023

ShadowDiffusion: When Degradation Prior Meets Diffusion Model for Shadow Removal

CVPR 2023poster

Recent deep learning methods have achieved promising results in image shadow removal. However, their restored images still suffer from unsatisfactory boundary artifacts, due to the lack of degradation prior and the deficiency in modeling capacity. Our work addresses these issues by proposing a unifi…

2023

ShadowFormer: Global Context Helps Shadow Removal

AAAI 2023technical

Recent deep learning methods have achieved promising results in image shadow removal. However, most of the existing approaches focus on working locally within shadow and non-shadow regions, resulting in severe artifacts around the shadow boundaries as well as inconsistent illumination between shadow…

2023

Unsupervised Deep Digital Staining for Microscopic Cell Images via Knowledge Distillation

ICASSP 2023accepted

Staining is critical to cell imaging and medical diagnosis, which is expensive, time-consuming, labor-intensive, and causes irreversible changes to cell tissues. Recent advances in deep learning enabled digital staining via supervised model training. However, it is difficult to obtain large-scale st…

Cited by 0SourceScholar
2023

WBCAtt: A White Blood Cell Dataset Annotated with Detailed Morphological Attributes

NeurIPS 2023poster

The examination of blood samples at a microscopic level plays a fundamental role in clinical diagnostics. For instance, an in-depth study of White Blood Cells (WBCs), a crucial component of our blood, is essential for diagnosing blood-related diseases such as leukemia and anemia. While multiple data…

2023

sRGB Real Noise Synthesizing With Neighboring Correlation-Aware Noise Model

CVPR 2023poster

Modeling and synthesizing real noise in the standard RGB (sRGB) domain is challenging due to the complicated noise distribution. While most of the deep noise generators proposed to synthesize sRGB real noise using an end-to-end trained model, the lack of explicit noise modeling degrades the quality…

2022

DVS-Voltmeter: Stochastic Process-Based Event Simulator for Dynamic Vision Sensors

ECCV 2022poster

"Recent advances in deep learning for event-driven applications with dynamic vision sensors (DVS) primarily rely on training over simulated data. However, most simulators ignore various physics-based characteristics of real DVS, such as the fidelity of event timestamps and comprehensive noise effect…

2022

Feature Augmentation Learning for Few-Shot Palmprint Image Recognition With Unconstrained Acquisition

ICASSP 2022accepted

Few-shot learning is challenging in unconstrained palmprint recognition, where the palmprint images are collected by unconstrained acquisitions, i.e., different imaging sensors, backgrounds, palm postures, and illumination conditions. Furthermore, due to the lack of unconstrained palmprint databases…

Cited by 0SourceScholar
2022

PIP: Physical Interaction Prediction via Mental Simulation with Span Selection

ECCV 2022poster

"Accurate prediction of physical interaction outcomes is a crucial component of human intelligence and is important for safe and efficient deployments of robots in the real world. While there are existing vision-based intuitive physics models that learn to predict physical interaction outcomes, they…

Cited by 7SourcePDFScholar
2022

Parameter-Free Style Projection for Arbitrary Image Style Transfer

ICASSP 2022accepted

Arbitrary image style transfer is a challenging task which aims to stylize a content image conditioned on arbitrary style images. In this task the feature-level content-style transformation plays a vital role for proper fusion of features. Existing feature transformation algorithms often suffer from…

Cited by 0SourceScholar
2022

Simultaneous Nonlocal Low-Rank And Deep Priors For Poisson Denoising

ICASSP 2022accepted

Poisson noise is a common electronic noise, which has widely occurred in various photo-limited imaging systems. However, due to signal-dependent and multiplicative characteristics for Poisson noise, Poisson denoising is still an open problem. In this paper, we propose a novel approach using simultan…

Cited by 0SourceScholar
2021

Learning Sparsifying Transforms for Image Reconstruction in Electrical Impedance Tomography

ICASSP 2021accepted

Electrical Impedance Tomography (EIT) is a fast and non-invasive imaging technology that reconstructs the internal electrical properties of a subject. However, its functionality is limited by low spatial resolution arising from an ill-posed and ill-conditioned inverse problem. Several sparsity-promo…

Cited by 0SourceScholar
2021

Recent Advances in Adversarial Training for Adversarial Robustness

IJCAI 2021poster

Adversarial training is one of the most effective approaches for deep learning models to defend against adversarial examples. Unlike other defense strategies, adversarial training aims to enhance the robustness of models intrinsically. During the past few years, adversarial training has been studi…

Cited by 596SourcePDFScholar
2021

Self-Convolution: A Highly-Efficient Operator for Non-Local Image Restoration

ICASSP 2021accepted

Constructing effective image priors is critical to solving ill-posed inverse problems, such as image restoration. Recent works proposed to exploit image non-local similarity for inverse problems by grouping similar patches, and demonstrated state-of-the-art results in many applications. However, com…

Cited by 0SourceScholar
2020

Generating Person Images with Appearance-aware Pose Stylizer

IJCAI 2020poster

Generation of high-quality person images is challenging, due to the sophisticated entanglements among image factors, e.g., appearance, pose, foreground, background, local details, global structures, etc. In this paper, we present a novel end-to-end framework to generate realistic person images based…

2018

Deepcasd: An End-to-End Approach for Multi-Spectral Image Super-Resolution

ICASSP 2018accepted

Multi-spectral (MS) image super-resolution aims to reconstruct super-resolved multi-channel images from their low-resolution images by regularizing the image to be reconstructed. Recently data-driven regularization techniques based on sparse modeling and deep learning have achieved substantial impro…

Cited by 0SourceScholar
2018

Non-Local Recurrent Network for Image Restoration

NeurIPS 2018poster

Many classic methods have shown non-local self-similarity in natural images to be an effective prior for image restoration. However, it remains unclear and challenging to make use of this intrinsic property via deep networks. In this paper, we propose a non-local recurrent network (NLRN) as the firs…

2017

Joint Adaptive Sparsity and Low-Rankness on the Fly: An Online Tensor Reconstruction Scheme for Video Denoising

ICCV 2017poster

Recent works on adaptive sparse and low-rank signal modeling have demonstrated their usefulness, especially in image/video processing applications. While a patch-based sparse model imposes local structure, low-rankness of the grouped patches exploits non-local correlation. Applying either approach a…

Cited by 56PDFScholar
2017

When sparsity meets low-rankness: Transform learning with non-local low-rank constraint for image restoration

ICASSP 2017accepted

Recent works on adaptive sparse signal modeling have demonstrated their usefulness in various image/video processing applications. As the popular synthesis dictionary learning methods involve NP-hard sparse coding and expensive learning steps, transform learning has recently received more interest f…

Cited by 0SourceScholar