← Search

Zhiwei Xiong

93 accepted papers

2026

Arbitrary-Shaped Image Generation via Spherical Neural Field Diffusion

ICLR 2026poster

Existing diffusion models excel at generating diverse content, but remain confined to fixed image shapes and lack the ability to flexibly control spatial attributes such as viewpoint, field-of-view (FOV), and resolution. To fill this gap, we propose Arbitrary-Shaped Image Generation (ASIG), the fir…

Cited by 0SourceScholar
2026

Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning

AAAI 2026technical

Large vision-language models (VLMs) for autonomous driving (AD) are evolving beyond perception and cognition tasks toward motion planning. However, we identify two critical challenges in this direction: (1) VLMs tend to learn shortcuts by relying heavily on history input information, achieving seemi

Cited by 0SourcePDFScholar
2026

Efficient Plug-and-Play Weight Refinement for Sparse Large Models

AAAI 2026technical

One-shot pruning efficiently compresses Large Language Models but produces coarse sparse weights, causing significant performance degradation. Traditional fine-tuning approaches to refine these weights are prohibitively expensive for large models. This highlights the need for a training-free weight

Cited by 0SourcePDFScholar
2026

FastGaMer: Efficient GainMap Learning for Practical Inverse Tone Mapping

CVPR 2026

Inverse tone mapping (ITM) becomes significantly harder when the SDR input is produced by local tone mapping, which jointly applies global radiometric compression and spatially varying adaptations that distort dynamic range, contrast, and channel-wise color ratios. Existing ITM methods ignore this d

Cited by 0SourceScholar
2026

Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search

CVPR 2026

Long video understanding presents significant challenges for vision-language models due to extremely long context windows. Existing solutions relying on naive chunking strategies with retrieval-augmented generation, typically suffer from information fragmentation and a loss of global coherence. We p

Cited by 0SourceScholar
2026

Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation

ICLR 2026poster

Lightweight 3D medical image segmentation remains constrained by a fundamental "efficiency / robustness conflict", particularly when processing complex anatomical structures and heterogeneous modalities. In this paper, we study how to redesign the framework based on the characteristics of high-dimen…

Cited by 0SourcecodeScholar
2026

RawMetaDiff: Unlocking Extreme Darkness from Dual-Exposure RAW with Meta-Guided Diffusion

CVPR 2026

Extreme low-light Raw image restoration remains challenging due to overwhelming noise and severe detail loss.In this paper, we exploit the potential of the dual-exposure setting for this severely ill-posed problem.Existing methods suffer from unreliable cross-exposure alignment, resulting in degrade

Cited by 0SourceScholar
2026

SHERPA: Fine-tuning Segment Anything Models with Task-relevant Guidance

ICML 2026poster

Segment Anything Models (SAMs) often struggle with certain specialized tasks. A common approach is to fine-tune models with specific task labels, but this often leads to overfitting, introduces model bias and significantly degrades their generalization ability. To overcome these challenges, we propo…

Cited by 0SourceScholar
2025

CBQ: Cross-Block Quantization for Large Language Models

ICLR 2025spotlight

Post-training quantization (PTQ) has played a pivotal role in compressing large language models (LLMs) at ultra-low costs. Although current PTQ methods have achieved promising results by addressing outliers and employing layer- or block-wise loss optimization techniques, they still suffer from signi…

Cited by 13SourcePDFScholar
2025

Event-Enhanced Blurry Video Super-Resolution

AAAI 2025technical

In this paper, we tackle the task of blurry video super-resolution (BVSR), aiming to generate high-resolution (HR) videos from low-resolution (LR) and blurry inputs. Current BVSR methods often fail to restore sharp details at high resolutions, resulting in noticeable artifacts and jitter due to insu…

2025

Event-boosted Deformable 3D Gaussians for Dynamic Scene Reconstruction

ICCV 2025poster

Deformable 3D Gaussian Splatting (3D-GS) is limited by missing intermediate motion information due to the low temporal resolution of RGB cameras. To address this, we introduce the first approach combining event cameras, which capture high-temporal-resolution, continuous motion data, with deformable…

2025

Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving

ICCV 2025poster

Existing benchmarks for Vision-Language Model (VLM) in autonomous driving (AD) primarily assess interpretability through open-form visual question answering (QA) within coarse-grained tasks, which remain insufficient to assess capabilities in complex driving scenarios. To this end, we introduce VLAD…

2025

GenFlow3D: Generative Scene Flow Estimation and Prediction on Point Cloud Sequences

ICCV 2025poster

Scene flow provides the fundamental information of the scene dynamics. Existing scene flow estimation methods typically rely on the correlation between only a consecutive point cloud pair, which makes them limited to the instantaneous state of the scene and face challenges in real-world scenarios wi…

2025

Generalizable Non-Line-of-Sight Imaging with Learnable Physical Priors

ICCV 2025poster

Non-line-of-sight (NLOS) imaging, recovering the hidden volume from indirect reflections, has attracted increasing attention due to its potential applications. Despite promising results, existing NLOS reconstruction approaches are constrained by the reliance on empirical physical priors, e.g., singl…

2025

HDR Image Generation via Gain Map Decomposed Diffusion

ICCV 2025poster

While diffusion models have demonstrated significant success in standard dynamic range (SDR) image synthesis, generating high dynamic range (HDR) images with higher luminance and broader color gamuts remains challenging. This arises primarily from two factors: (1) The incompatibility between pretrai…

2025

Learning Gain Map for Inverse Tone Mapping

ICLR 2025poster

For a more compatible and consistent high dynamic range (HDR) viewing experience, a new image format with a double-layer structure has been developed recently, which incorporates an auxiliary Gain Map (GM) within a standard dynamic range (SDR) image for adaptive HDR display. This new format motivate…

2025

MaskTwins: Dual-form Complementary Masking for Domain-Adaptive Image Segmentation

ICML 2025poster

Recent works have correlated Masked Image Modeling (MIM) with consistency regularization in Unsupervised Domain Adaptation (UDA). However, they merely treat masking as a special form of deformation on the input images and neglect the theoretical analysis, which leads to a superficial understanding o…

2025

Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding

ICML 2025poster

Achieving high-fidelity audio compression while preserving perceptual quality across diverse audio types remains a significant challenge in Neural Audio Coding (NAC). This paper introduces MUFFIN, a fully convolutional NAC framework that leverages psychoacoustically guided multi-band frequency recon…

2025

S2D-LFE: Sparse-to-Dense Light Field Event Generation

CVPR 2025poster

In this paper, we present S2D-LFE, an innovative approach for sparse-to-dense light field event generation. For the first time to our knowledge, S2D-LFE enables controllable novel view synthesis only from sparse-view light field event (LFE) data, and addresses three critical challenges for the LFE g…

2025

TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction

ICCV 2025poster

We present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQVAE) and a Generative Pre-trained Transformer (GPT) to generate high-quality 3D assets. The core insight of this work is to migrate the multimodal unification and promising learning capabiliti…

2024

Cross-Dimension Affinity Distillation for 3D EM Neuron Segmentation

CVPR 2024poster

Accurate 3D neuron segmentation from electron microscopy (EM) volumes is crucial for neuroscience research. However the complex neuron morphology often leads to over-merge and over-segmentation results. Recent advancements utilize 3D CNNs to predict a 3D affinity map with improved accuracy but suffe…

2024

Event-assisted Low-Light Video Object Segmentation

CVPR 2024poster

In the realm of video object segmentation (VOS) the challenge of operating under low-light conditions persists resulting in notably degraded image quality and compromised accuracy when comparing query and memory frames for similarity computation. Event cameras characterized by their high dynamic ran…

2024

Exploiting Dual-Correlation for Multi-frame Time-of-Flight Denoising

ECCV 2024poster

"Recent advancements in Time-of-Flight (ToF) depth denoising have achieved impressive results in removing Multi-Path Interference (MPI) and shot noise. However, existing methods only utilize a single frame of ToF data, neglecting the correlation between frames. In this paper, we propose the first le…

2024

Learning Large-Factor EM Image Super-Resolution with Generative Priors

CVPR 2024poster

As the mainstream technique for capturing images of biological specimens at nanometer resolution electron microscopy (EM) is extremely time-consuming for scanning wide field-of-view (FOV) specimens. In this paper we investigate a challenging task of large-factor EM image super-resolution (EMSR) whic…

2024

Learning Multimodal Volumetric Features for Large-Scale Neuron Tracing

AAAI 2024technical

The current neuron reconstruction pipeline for electron microscopy (EM) data usually includes automatic image segmentation followed by extensive human expert proofreading. In this work, we aim to reduce human workload by predicting connectivity between over-segmented neuron pieces, taking both micro…

2024

Learning Multiscale Consistency for Self-Supervised Electron Microscopy Instance Segmentation

ICASSP 2024accepted

Electron microscopy (EM) images are notoriously challenging to segment due to their complex structures and lack of effective annotations. Fortunately, large-scale self-supervised pretraining offers a promising solution by allowing us to acquire prior knowledge of cell and subcellular tissue structur…

Cited by 0SourceScholar
2024

Test-Time Adaptation via Style and Structure Guidance for Histological Image Registration

AAAI 2024technical

Image registration plays a crucial role in histological image analysis, encompassing tasks like multi-modality fusion and disease grading. Traditional registration methods optimize objective functions for each image pair, yielding reliable accuracy but demanding heavy inference burdens. Recently, l…

Cited by 1SourcePDFScholar
2024

Toward Dynamic Non-Line-of-Sight Imaging with Mamba Enforced Temporal Consistency

NeurIPS 2024poster

Dynamic reconstruction in confocal non-line-of-sight imaging encounters great challenges since the dense raster-scanning manner limits the practical frame rate. A fewer pioneer works reconstruct high-resolution volumes from the under-scanning transient measurements but overlook temporal consistency…

2024

Towards Generalizable Tumor Synthesis

CVPR 2024poster

Tumor synthesis enables the creation of artificial tumors in medical images facilitating the training of AI models for tumor detection and segmentation. However success in tumor synthesis hinges on creating visually realistic tumors that are generalizable across multiple organs and furthermore the r…

2023

A Soma Segmentation Benchmark in Full Adult Fly Brain

CVPR 2023poster

Neuron reconstruction in a full adult fly brain from high-resolution electron microscopy (EM) data is regarded as a cornerstone for neuroscientists to explore how neurons inspire intelligence. As the central part of neurons, somas in the full brain indicate the origin of neurogenesis and neural func…

2023

Adaptive Template Transformer for Mitochondria Segmentation in Electron Microscopy Images

ICCV 2023poster

Mitochondria, as tiny structures within the cell, are of significant importance to study cell functions for biological and clinical analysis. And exploring how to automatically segment mitochondria in electron microscopy (EM) images has attracted increasing attention. However, most of existing metho…

Cited by 19PDFScholar
2023

Appearance Prompt Vision Transformer for Connectome Reconstruction

IJCAI 2023poster

Neural connectivity reconstruction aims to understand the function of biological reconstruction and promote basic scientific research. The intricate morphology and densely intertwined branches make it an extremely challenging task. Most previous best-performing methods adopt affinity learning or met…

Cited by 16SourcePDFScholar
2023

Camouflaged Instance Segmentation via Explicit De-Camouflaging

CVPR 2023highlight

Camouflaged Instance Segmentation (CIS) aims at predicting the instance-level masks of camouflaged objects, which are usually the animals in the wild adapting their appearance to match the surroundings. Previous instance segmentation methods perform poorly on this task as they are easily disturbed b…

Cited by 37SourcePDFScholar
2023

CutMIB: Boosting Light Field Super-Resolution via Multi-View Image Blending

CVPR 2023poster

Data augmentation (DA) is an efficient strategy for improving the performance of deep neural networks. Recent DA strategies have demonstrated utility in single image super-resolution (SR). Little research has, however, focused on the DA strategy for light field SR, in which multi-view information ut…

2023

Deep Non-line-of-sight Imaging from Under-scanning Measurements

NeurIPS 2023poster

Active confocal non-line-of-sight (NLOS) imaging has successfully enabled seeing around corners relying on high-quality transient measurements. However, acquiring spatial-dense transient measurement is time-consuming, raising the question of how to reconstruct satisfactory results from under-scannin…

2023

Depth Estimation From Indoor Panoramas With Neural Scene Representation

CVPR 2023poster

Depth estimation from indoor panoramas is challenging due to the equirectangular distortions of panoramas and inaccurate matching. In this paper, we propose a practical framework to improve the accuracy and efficiency of depth estimation from multi-view indoor panoramic images with the Neural Radian…

2023

DualRel: Semi-Supervised Mitochondria Segmentation From a Prototype Perspective

CVPR 2023poster

Automatic mitochondria segmentation enjoys great popularity with the development of deep learning. However, existing methods rely heavily on the labor-intensive manual gathering by experienced domain experts. And naively applying semi-supervised segmentation methods in the natural image field to mit…

Cited by 25SourcePDFScholar
2023

GET: Group Event Transformer for Event-Based Vision

ICCV 2023poster

Event cameras are a type of novel neuromorphic sen-sor that has been gaining increasing attention. Existing event-based backbones mainly rely on image-based designs to extract spatial information within the image transformed from events, overlooking important event properties like time and polarity.…

Cited by 107PDFcodeScholar
2023

Generalized Lightness Adaptation with Channel Selective Normalization

ICCV 2023poster

Lightness adaptation is vital to the success of image processing to avoid unexpected visual deterioration, which covers multiple aspects, e.g., low-light image enhancement, image retouching, and inverse tone mapping. Existing methods typically work well on their trained lightness conditions but perf…

Cited by 20PDFcodeScholar
2023

Hierarchical Prompt Learning for Multi-Task Learning

CVPR 2023poster

Vision-language models (VLMs) can effectively transfer to various vision tasks via prompt learning. Real-world scenarios often require adapting a model to multiple similar yet distinct tasks. Existing methods focus on learning a specific prompt for each task, limiting the ability to exploit potentia…

Cited by 39SourcePDFScholar
2023

Learning Cross-Representation Affinity Consistency for Sparsely Supervised Biomedical Instance Segmentation

ICCV 2023poster

Sparse instance-level supervision has recently been explored to address insufficient annotation in biomedical instance segmentation, which is easier to annotate crowded instances and better preserves instance completeness for 3D volumetric datasets compared to common semi-supervision.In this paper,…

Cited by 8PDFcodeScholar
2023

Learning Sample Relationship for Exposure Correction

CVPR 2023poster

Exposure correction task aims to correct the underexposure and its adverse overexposure images to the normal exposure in a single network. As well recognized, the optimization flow is opposite. Despite the great advancement, existing exposure correction methods are usually trained with a mini-batch…

Cited by 49SourcePDFScholar
2023

Learning Steerable Function for Efficient Image Resampling

CVPR 2023poster

Image resampling is a basic technique that is widely employed in daily applications. Existing deep neural networks (DNNs) have made impressive progress in resampling performance. Yet these methods are still not the perfect substitute for interpolation, due to the issues of efficiency and continuous…

Cited by 11SourcePDFScholar
2023

NLOST: Non-Line-of-Sight Imaging With Transformer

CVPR 2023poster

Time-resolved non-line-of-sight (NLOS) imaging is based on the multi-bounce indirect reflections from the hidden objects for 3D sensing. Reconstruction from NLOS measurements remains challenging especially for complicated scenes. To boost the performance, we present NLOST, the first transformer-base…

Cited by 29SourcePDFScholar
2023

Progressive Spatio-Temporal Alignment for Efficient Event-Based Motion Estimation

CVPR 2023poster

In this paper, we propose an efficient event-based motion estimation framework for various motion models. Different from previous works, we design a progressive event-to-map alignment scheme and utilize the spatio-temporal correlations to align events. In detail, we progressively align sampled event…

2023

Self-Supervised Neuron Segmentation with Multi-Agent Reinforcement Learning

IJCAI 2023poster

The performance of existing supervised neuron segmentation methods is highly dependent on the number of accurate annotations, especially when applied to large scale electron microscopy (EM) data. By extracting semantic information from unlabeled data, self-supervised methods can improve the performa…

2023

Style Projected Clustering for Domain Generalized Semantic Segmentation

CVPR 2023poster

Existing semantic segmentation methods improve generalization capability, by regularizing various images to a canonical feature space. While this process contributes to generalization, it weakens the representation inevitably. In contrast to existing methods, we instead utilize the difference betwee…

Cited by 40SourcePDFScholar
2023

Toward RAW Object Detection: A New Benchmark and a New Model

CVPR 2023poster

In many computer vision applications (e.g., robotics and autonomous driving), high dynamic range (HDR) data is necessary for object detection algorithms to handle a variety of lighting conditions, such as strong glare. In this paper, we aim to achieve object detection on RAW sensor data, which natur…

Cited by 28SourcePDFScholar
2023

Transition-constant Normalization for Image Enhancement

NeurIPS 2023spotlight

Normalization techniques that capture image style by statistical representation have become a popular component in deep neural networks. Although image enhancement can be considered as a form of style transformation, there has been little exploration of how normalization affect the enhancement perfo…

2023

Why Is the Winner the Best?

CVPR 2023poster

International benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and…

Cited by 29SourcePDFScholar
2022

Biological Instance Segmentation with a Superpixel-Guided Graph

IJCAI 2022poster

Recent advanced proposal-free instance segmentation methods have made significant progress in biological images. However, existing methods are vulnerable to local imaging artifacts and similar object appearances, resulting in over-merge and over-segmentation. To reduce these two kinds of errors, we…

2022

Deep Fourier-Based Exposure Correction Network with Spatial-Frequency Interaction

ECCV 2022poster

"Images captured under incorrect exposures unavoidably suffer from mixed degradations of lightness and structures. Most existing deep learning-based exposure correction methods separately restore such degradations in the spatial domain. In this paper, we present a new perspective for exposure correc…

2022

Degradation-Agnostic Correspondence From Resolution-Asymmetric Stereo

CVPR 2022poster

In this paper, we study the problem of stereo matching from a pair of images with different resolutions, e.g., those acquired with a tele-wide camera system. Due to the difficulty of obtaining ground-truth disparity labels in diverse real-world systems, we start from an unsupervised learning perspec…

Cited by 10PDFScholar
2022

Efficient Model-Driven Network for Shadow Removal

AAAI 2022technical

Deep Convolutional Neural Networks (CNNs) based methods have achieved significant breakthroughs in the task of single image shadow removal. However, the performance of these methods remains limited for several reasons. First, the existing shadow illumination model ignores the spatially variant prope…

2022

Exploiting Rigidity Constraints for LiDAR Scene Flow Estimation

CVPR 2022poster

Previous LiDAR scene flow estimation methods, especially recurrent neural networks, usually suffer from structure distortion in challenging cases, such as sparse reflection and motion occlusions. In this paper, we propose a novel optimization method based on a recurrent neural network to predict LiD…

Cited by 40PDFScholar
2022

Exposure Normalization and Compensation for Multiple-Exposure Correction

CVPR 2022poster

Images captured with improper exposures usually bring unsatisfactory visual effects. Previous works mainly focus on either underexposure or overexposure correction, resulting in poor generalization to various exposures. An alternative solution is to mix the multiple exposure data for training a sing…

Cited by 60PDFScholar
2022

Learning to Model Pixel-Embedded Affinity for Homogeneous Instance Segmentation

AAAI 2022technical

Homogeneous instance segmentation aims to identify each instance in an image where all interested instances belong to the same category, such as plant leaves and microscopic cells. Recently, proposal-free methods, which straightforwardly generate instance-aware information to group pixels into diffe…

2022

MuLUT: Cooperating Multiple Look-Up Tables for Efficient Image Super-Resolution

ECCV 2022poster

"The high-resolution screen of edge devices stimulates a strong demand for efficient image super-resolution (SR). An emerging research, SR-LUT, responds to this demand by marrying the look-up table (LUT) with learning-based SR methods. However, the size of a single LUT grows exponentially with the i…

Cited by 39SourcePDFScholar
2022

Quantization-Aware Deep Optics for Diffractive Snapshot Hyperspectral Imaging

CVPR 2022poster

Diffractive snapshot hyperspectral imaging based on the deep optics framework has been striving to capture the spectral images of dynamic scenes. However, existing deep optics frameworks all suffer from the mismatch between the optical hardware and the reconstruction algorithm due to the quantizatio…

Cited by 47PDFScholar
2022

Recurrent Dynamic Embedding for Video Object Segmentation

CVPR 2022poster

Space-time memory (STM) based video object segmentation (VOS) networks usually keep increasing memory bank every several frames, which shows excellent performance. However, 1) the hardware cannot withstand the ever-increasing memory requirements as the video length increases. 2) Storing lots of info…

Cited by 95PDFcodeScholar
2022

Retriever: Learning Content-Style Representation as a Token-Level Bipartite Graph

ICLR 2022poster

This paper addresses the unsupervised learning of content-style decomposed representation. We first give a definition of style and then model the content-style representation as a token-level bipartite graph. An unsupervised framework, named Retriever, is proposed to learn such representations. Firs…

2022

Towards Real-World HDRTV Reconstruction: A Data Synthesis-Based Approach

ECCV 2022poster

"Existing deep learning based HDRTV reconstruction methods assume one kind of tone mapping operators (TMOs) as the degradation procedure to synthesize SDRTV-HDRTV pairs for supervised training. In this paper, we argue that, although traditional TMOs exploit efficient dynamic range compression priors…

2021

Training Spiking Neural Networks with Accumulated Spiking Flow

AAAI 2021technical

The fast development of neuromorphic hardwares promotes Spiking Neural Networks (SNNs) to a thrilling research avenue. Current SNNs, though much efficient, are less effective compared with leading Artificial Neural Networks (ANNs) especially in supervised learning tasks. Recent efforts further demon…

2021

Unfolding Taylor's Approximations for Image Restoration

NeurIPS 2021poster

Deep learning provides a new avenue for image restoration, which demands a delicate balance between fine-grained details and high-level contextualized information during recovering the latent clear image. In practice, however, existing methods empirically construct encapsulated end-to-end mapping ne…

Cited by 26SourcePDFScholar
2021

Unsupervised Visual Representation Learning by Tracking Patches in Video

CVPR 2021poster

Inspired by the fact that human eyes continue to develop tracking ability in early and middle childhood, we propose to use tracking as a proxy task for a computer vision system to learn the visual representations. Modelled on the Catch game played by the children, we design a Catch-the-Patch (CtP) g…

Cited by 31PDFcodeScholar
2020

Photon-Efficient 3D Imaging with A Non-Local Neural Network

ECCV 2020poster

Photon-efficient imaging has enabled a number of applications relying on single-photon sensors that can capture a 3D image with as few as one photon per pixel. In practice, however, measurements of low photon counts are often mixed with heavy background noise, which poses a great challenge for exist…

2020

Spatial Hierarchy Aware Residual Pyramid Network for Time-of-Flight Depth Denoising

ECCV 2020poster

Time-of-Flight (ToF) sensors have been increasingly used on mobile devices for depth sensing. However, the existence of noise, such as Multi-Path Interference (MPI) and shot noise, degrades the ToF imaging quality. Previous CNN-based methods remove ToF depth noise without considering the spatial hie…

2019

SPM-Tracker: Series-Parallel Matching for Real-Time Visual Object Tracking

CVPR 2019poster

The greatest challenge facing visual object tracking is the simultaneous requirements on robustness and discrimination power. In this paper, we propose a SiamFC-based tracker, named SPM-Tracker, to tackle this challenge. The basic idea is to address the two requirements in two separate matching stag…

Cited by 290PDFScholar
2015

High-Speed Hyperspectral Video Acquisition With a Dual-Camera Architecture

CVPR 2015poster

We propose a novel dual-camera design to acquire 4D high-speed hyperspectral (HSHS) videos with high spatial and spectral resolution. Our work has two key technical contributions. First, we build a dual-camera system that simultaneously captures a panchromatic video at a high frame rate and a hypers…

Cited by 116SourcePDFScholar