← Search

Xiaopeng Fan

35 accepted papers

2026

Beyond Single Solution: Multi-Hypothesis Deep Unfolding Network for Image Compressive Sensing

CVPR 2026

Recent deep unfolding networks (DUNs) have advanced Compressive Sensing (CS) by effectively integrating iterative optimization with deep learning architectures. However, most CS approaches predominantly confine their inference to a single solution space, neglecting the inherent ill-posedness of CS p

Cited by 0SourceScholar
2026

Conflict-Aware Additive Guidance for Flow Models under Compositional Rewards

ICML 2026poster

Inference-time guided sampling steers state-of-the-art diffusion and flow models without fine-tuning by interpreting the generation process as a controllable trajectory. This provides a simple and flexible way to inject external constraints (e.g., cost functions or pre-trained verifiers) for control…

Cited by 0SourceScholar
2026

Hyperbolic Hierarchical Alignment Reasoning Network for Text-3D Retrieval

AAAI 2026technical

With the daily influx of 3D data on the internet, text-3D retrieval has gained increasing attention. However, current methods face two major challenges: Hierarchy Representation Collapse (HRC) and Redundancy-Induced Saliency Dilution (RISD). HRC compresses abstract-to-specific and whole-to-part hier

Cited by 0SourcePDFScholar
2026

Perceptual Quality Assessment of 3D Gaussian Splatting: A Subjective Dataset and Prediction Metric

AAAI 2026technical

With the rapid advancement of 3D visualization, 3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time, high-fidelity rendering. While prior research has emphasized algorithmic performance and visual fidelity, the perceptual quality of 3DGS-rendered content, especially under v

Cited by 0SourcePDFScholar
2026

Spk2VidNet: A Hierarchical Recurrent Architecture for High-Fidelity Video Reconstruction from Long Spike-Camera Streams

CVPR 2026

Spike camera is a neuromorphic vision sensor with ultra-high temporal resolution, capable of capturing fast-moving scenes by firing a stream of binary spikes. However, its relatively low spatial resolution limits the acquisition of fine-grained visual details, motivating research on spike camera sup

Cited by 0SourceScholar
2026

T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low Bitrates

AAAI 2026technical

Recent advances in video generation techniques have given rise to an emerging paradigm of generative video coding for Ultra-Low Bitrate (ULB) scenarios by leveraging powerful generative priors. However, most existing methods are limited by domain specificity (e.g., facial or human videos) or excessi

Cited by 0SourcePDFScholar
2025

Asynchronous Collaborative Graph Representation for Frames and Events

CVPR 2025poster

Integrating frames and events has become a widely accepted solution for various tasks in challenging scenarios. However, most multimodal methods directly convert events into image-like formats synchronized with frames and process each stream through separate two-branch backbones, making it difficult…

2025

CASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional Video

CVPR 2025poster

Omnidirectional videos (ODVs) present distinct challenges for accurate audio-visual saliency prediction due to their immersive nature, which combines spatial audio with panoramic visuals to enhance the user experience. While auditory cues are crucial for guiding visual attention across the panoramic…

Cited by 0SourcePDFScholar
2025

DexFlyWheel: A Scalable and Self-improving Data Generation Framework for Dexterous Manipulation

NeurIPS 2025spotlight

Dexterous manipulation is critical for advancing robot capabilities in real-world applications, yet diverse and high-quality datasets remain scarce. Existing data collection methods either rely on human teleoperation or require significant human engineering, or generate data with limited diversity,…

Cited by 0SourceScholar
2025

Digging into Intrinsic Contextual Information for High-fidelity 3D Point Cloud Completion

AAAI 2025technical

The common occurrence of occlusion-induced incompleteness in point clouds has made point cloud completion (PCC) a highly-concerned task in the field of geometric processing. Existing PCC methods typically produce complete point clouds from partial point clouds in a coarse-to-fine paradigm, with the…

2025

Hyperbolic-Constraint Point Cloud Reconstruction from Single RGB-D Images

AAAI 2025technical

Reconstructing desired objects and scenes has long been a primary goal in 3D computer vision. Single-view point cloud reconstruction has become a popular technique due to its low cost and accurate results. However, single-view reconstruction methods often rely on expensive CAD models and complex geo…

Cited by 0SourcePDFScholar
2025

ISP2HRNet: Learning to Reconstruct High Resolution Image from Irregularly Sampled Pixels via Hierarchical Gradient Learning

ICCV 2025poster

While image signals are typically defined on a regular 2D grid, there are scenarios where they are only available at irregular positions. In such cases, reconstructing a complete image on regular grid is essential. This paper introduces ISP2HRNet, an end-to-end network designed to reconstruct high r…

2025

Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval

AAAI 2025technical

Due to the challenges in acquiring paired Text-3D data and the inherent irregularity of 3D data structures, combined representation learning of 3D point clouds and text remains unexplored. In this paper, we propose a novel Riemann-based Multi-scale Attention Reasoning Network (RMARN) for text-3D ret…

2025

SAMPLE: Semantic Alignment through Temporal-Adaptive Multimodal Prompt Learning for Event-Based Open-Vocabulary Action Recognition

ICCV 2025poster

Open-vocabulary action recognition (OVAR) extends recognition systems to identify unseen action categories. While large-scale vision-language models (VLMs) like CLIP have enabled OVAR in image domains, their adaptation to event data remains underexplored. Event cameras offer high temporal resolution…

2025

Self-Supervised Learning for Color Spike Camera Reconstruction

CVPR 2025poster

Spike camera is a kind of neuromorphic camera with ultra-high temporal resolution, which can capture dynamic scenes by continuously firing spike signals. To capture color information, a color filter array (CFA) is employed on the sensor of the spike camera, resulting in Bayer-pattern spike streams.…

2025

Spatial Frequency Interleaving Residual Autoencoder for Indoor Radio Map Reconstruction

ICASSP 2025accepted

Indoor radio maps with frequency domain data are difficult to reconstruct when only limited measurements at a few locations are available. Naive convolutional neural networks suffer from flawed structures in the frequency domain when predicting these radio maps, resulting in overly smoothed predicti…

Cited by 0SourceScholar
2025

Spk2SRImgNet: Super-Resolve Dynamic Scene from Spike Stream via Motion Aligned Collaborative Filtering

CVPR 2025poster

Spike camera is a kind of neuromorphic camera that records dynamic scenes by firing a stream of binary spikes with extremely high temporal resolution. It demonstrates great potential for vision tasks in high-speed scenarios. One limitation in its current implementation is the relatively low spatial…

Cited by 0SourcePDFScholar
2025

TS-Net: Assembling Task-specific Features from Multiple Feature Levels for Multi-task Learning

ICASSP 2025accepted

Multi-task learning (MTL) has become an attractive topic that leverages shared knowledge to improve performance and enhance generalization. However, most existing works neglect the varying contribution of multi-level features to sub-task representations. In this paper, we explore the impact of multi…

Cited by 0SourceScholar
2025

Text-Guided Editable 3D City Scene Generation

ICASSP 2025accepted

The automated generation of 3D city scenes has attracted considerable attention due to its broad applications in areas such as virtual reality, urban planning, and digital media. Traditional approaches for constructing 3D city environments typically depend on labor-intensive manual modeling or the u…

Cited by 0SourceScholar
2025

Training-Free Task Planning by Parsing Language Signals With Common Sense

ICASSP 2025accepted

Task planning refers to autonomously organizing actions in response to instruction signals, especially language signals. Previous reinforcement learning and imitation learning methods always require a large amount of task-related data (data interacting with the environment or expert demonstrations)…

Cited by 0SourceScholar
2024

Boosting Spike Camera Image Reconstruction from a Perspective of Dealing with Spike Fluctuations

CVPR 2024poster

As a bio-inspired vision sensor with ultra-high speed spike cameras exhibit great potential in recording dynamic scenes with high-speed motion or drastic light changes. Different from traditional cameras each pixel in spike cameras records the arrival of photons continuously by firing binary spikes…

2024

Hyper-MD: Mesh Denoising with Customized Parameters Aware of Noise Intensity and Geometric Characteristics

CVPR 2024poster

Mesh denoising (MD) is a critical task in geometry processing as meshes from scanning or AIGC techniques are susceptible to noise contamination. The challenge of MD lies in the diverse nature of mesh facets in terms of geometric characteristics and noise distributions. Despite recent advancements in…

Cited by 0SourcePDFScholar
2024

Joint Demosaicing and Denoising for Spike Camera

AAAI 2024technical

As a neuromorphic camera with high temporal resolution, spike camera can capture dynamic scenes with high-speed motion. Recently, spike camera with a color filter array (CFA) has been developed for color imaging. There are some methods for spike camera demosaicing to reconstruct color images from Ba…

2024

QKFormer: Hierarchical Spiking Transformer using Q-K Attention

NeurIPS 2024spotlight

Spiking Transformers, which integrate Spiking Neural Networks (SNNs) with Transformer architectures, have attracted significant attention due to their potential for low energy consumption and high performance. However, there remains a substantial gap in performance between SNNs and Artificial Neural…

2024

Super-Resolution Reconstruction from Bayer-Pattern Spike Streams

CVPR 2024poster

Spike camera is a neuromorphic vision sensor that can capture highly dynamic scenes by generating a continuous stream of binary spikes to represent the arrival of photons at very high temporal resolution. Equipped with Bayer color filter array (CFA) color spike camera (CSC) has been invented to capt…

2024

Toward a Stable, Fair, and Comprehensive Evaluation of Object Hallucination in Large Vision-Language Models

NeurIPS 2024poster

Given different instructions, large vision-language models (LVLMs) exhibit different degrees of object hallucinations, posing a significant challenge to the evaluation of object hallucinations. Overcoming this challenge, existing object hallucination evaluation methods average the results obtained f…

Cited by 4SourcePDFScholar
2024

Virtual Immunohistochemistry Staining for Histological Images Assisted by Weakly-supervised Learning

CVPR 2024poster

Recently virtual staining technology has greatly promoted the advancement of histopathology. Despite the practical successes achieved the outstanding performance of most virtual staining methods relies on hard-to-obtain paired images in training. In this paper we propose a method for virtual immunoh…

2023

Collaborative Audio-Visual Event Localization Based on Sequential Decision and Cross-Modal Consistency

ICASSP 2023accepted

We focus on the audio-visual event (AVE) localization task, which refers to locating the segments with AVE and identifying their event categories. Since different event-relevant video segments often describe different aspects of an AVE, they can complement each other. However, current approaches mod…

Cited by 0SourceScholar
2023

Interactive Object Placement with Reinforcement Learning

ICML 2023poster

Object placement aims to insert a foreground object into a background image with a suitable location and size to create a natural composition. To predict a diverse distribution of placements, existing methods usually establish a one-to-one mapping from random vectors to the placements. However, thes…

Cited by 6SourcePDFScholar
2022

Distributed Audio-Visual Parsing Based On Multimodal Transformer and Deep Joint Source Channel Coding

ICASSP 2022accepted

Audio-visual parsing (AVP) is a newly emerged multimodal perception task, which detects and classifies audio-visual events in video. However, most existing AVP networks only use a simple attention mechanism to guide audio-visual multimodal events, and are implemented in a single end. This makes it u…

Cited by 0SourceScholar
2022

Learning Optical Flow from Continuous Spike Streams

NeurIPS 2022accept

Spike camera is an emerging bio-inspired vision sensor with ultra-high temporal resolution. It records scenes by accumulating photons and outputting continuous binary spike streams. Optical flow is a key task for spike cameras and their applications. A previous attempt has been made for spike-based…

2021

SNR-Adaptive Deep Joint Source-Channel Coding for Wireless Image Transmission

ICASSP 2021accepted

Considering the problem of joint source-channel coding (JSCC) for multi-user transmission of images over noisy channels, an autoencoder-based novel deep joint source-channel coding scheme is proposed in this paper. In the proposed JSCC scheme, the decoder can estimate the signal-to-noise ratio (SNR)…

Cited by 0SourceScholar
2019

Multiscale Directional Fusion for Depth Map Super Resolution with Denoising

ICASSP 2019accepted

To tackle three main problems in depth map super resolution (SR) process, which are texture copy artifacts, blurred edge artifacts and jagged edge artifacts, we propose a depth map super resolution with denoising method based on multiscale directional fusion via nonsubsampled contourlet transform (N…

Cited by 0SourceScholar