← Search

Zhan Ma

31 accepted papers

2026

DiT-IC: Aligned Diffusion Transformer for Efficient Image Compression

CVPR 2026

Diffusion-based image compression has recently shown outstanding perceptual fidelity, yet its practicality is hindered by prohibitive sampling overhead and high memory usage.Most existing diffusion codecs employ UNet architectures, where hierarchical downsampling forces diffusion to operate in shall

Cited by 0SourcecodeScholar
2026

Reinforced Rate Control for Neural Video Compression via Inter-Frame Rate–Distortion Awareness

AAAI 2026technical

Neural video compression (NVC) has demonstrated superior compression efficiency, yet effective rate control remains a significant challenge due to complex temporal dependencies. Existing rate control schemes typically leverage frame content to capture distortion interactions, overlooking inter-frame

Cited by 0SourcePDFScholar
2026

Taming Hierarchical Image Coding Optimization: A Spectral Regularization Perspective

ICLR 2026poster

Hierarchical coding offers distinct advantages for learned image compression by capturing multi-scale representations to support scale-wise modeling and enable flexible quality scalability, making it a promising alternative to single-scale models. However, its practical performance remains limited.…

Cited by 0SourceScholar
2025

Depth-Guided Bundle Sampling for Efficient Generalizable Neural Radiance Field Reconstruction

CVPR 2025poster

Recent advancements in generalizable novel view synthesis have achieved impressive quality through interpolation between nearby views. However, rendering high-resolution images remains computationally intensive due to the need for dense sampling of all rays. Recognizing that natural scenes are typic…

2025

GoLF-NRT: Integrating Global Context and Local Geometry for Few-Shot View Synthesis

CVPR 2025poster

Neural Radiance Fields (NeRF) have transformed novel view synthesis by modeling scene-specific volumetric representations directly from images. While generalizable NeRF models can generate novel views across unknown scenes by learning latent ray representations, their performance heavily depends on…

2025

Integrating Adaptive Sampling for Optimal Learned Video Compression

ICASSP 2025accepted

We propose a novel adaptive prediction network that dynamically determines the optimal sampling factor and Lagrangian multiplier for encoding each frame, guided by sequential information. By exploiting spatio-temporal redundancy through adaptive sampling, our method reduces bitrate consumption while…

Cited by 0SourceScholar
2025

Mitigating Ambiguities in 3D Classification with Gaussian Splatting

CVPR 2025poster

3D classification with point cloud input is a fundamental problem in 3D vision. However, due to the discrete nature and the insufficient material description of point cloud representations, there are ambiguities in distinguishing wire-like and flat surfaces, as well as transparent or reflective obje…

Cited by 0SourcePDFScholar
2025

Neural B-frame Video Compression with Bi-directional Reference Harmonization

NeurIPS 2025poster

Neural video compression (NVC) has made significant progress in recent years, while neural B-frame video compression (NBVC) remains underexplored compared to P-frame compression. NBVC can adopt bi-directional reference frames for better compression performance. However, NBVC's hierarchical coding ma…

Cited by 0SourcecodeScholar
2025

On Quantizing Neural Representation for Variable-Rate Video Coding

ICLR 2025spotlight

This work introduces NeuroQuant, a novel post-training quantization (PTQ) approach tailored to non-generalized Implicit Neural Representations for variable-rate Video Coding (INR-VC). Unlike existing methods that require extensive weight retraining for each target bitrate, we hypothesize that variab…

2025

RENO: Real-Time Neural Compression for 3D LiDAR Point Clouds

CVPR 2025poster

Despite the substantial advancements demonstrated by learning-based neural models in the LiDAR Point Cloud Compression (LPCC) task, realizing real-time compression--an indispensable criterion for numerous industrial applications--remains a formidable challenge. This paper proposes RENO, the first re…

2025

Towards Loss-Resilient Image Coding for Unstable Satellite Networks

AAAI 2025technical

Geostationary Earth Orbit (GEO) satellite communication demonstrates significant advantages in emergency short burst data services. However, unstable satellite networks, particularly those with frequent packet loss, present a severe challenge to accurate image transmission. To address it, we propose…

2025

Ultra Lowrate Image Compression with Semantic Residual Coding and Compression-aware Diffusion

ICML 2025poster

Existing multimodal large model-based image compression frameworks often rely on a fragmented integration of semantic retrieval, latent compression, and generative models, resulting in suboptimal performance in both reconstruction fidelity and coding efficiency. To address these challenges, we propo…

Cited by 0SourcePDFScholar
2024

All-in-One Image Coding for Joint Human-Machine Vision with Multi-Path Aggregation

NeurIPS 2024poster

Image coding for multi-task applications, catering to both human perception and machine vision, has been extensively investigated. Existing methods often rely on multiple task-specific encoder-decoder pairs, leading to high overhead of parameter and bitrate usage, or face challenges in multi-objecti…

2024

Another Way to the Top: Exploit Contextual Clustering in Learned Image Coding

AAAI 2024technical

While convolution and self-attention are extensively used in learned image compression (LIC) for transform coding, this paper proposes an alternative called Contextual Clustering based LIC (CLIC) which primarily relies on clustering operations and local attention for correlation characterization and…

Cited by 8SourcePDFScholar
2024

EmoTalk3D: High-Fidelity Free-View Synthesis of Emotional 3D Talking Head

ECCV 2024poster

"We present a novel approach for synthesizing 3D talking heads with controllable emotion, featuring enhanced lip synchronization and rendering quality. Despite significant progress in the field, prior methods still suffer from multi-view consistency and a lack of emotional expressiveness. To address…

2024

Encoding Auxiliary Information to Restore Compressed Point Cloud Geometry

IJCAI 2024poster

The standardized Geometry-based Point Cloud Compression (G-PCC) suffers from limited coding performance and low-quality reconstruction. To address this, we propose AuxGR, a performance-complexity tradeoff solution for point cloud geometry restoration: leveraging auxiliary bitstream to enhance the qu…

2024

FINER: Flexible Spectral-bias Tuning in Implicit NEural Representation by Variable-periodic Activation Functions

CVPR 2024poster

Implicit Neural Representation (INR) which utilizes a neural network to map coordinate inputs to corresponding attributes is causing a revolution in the field of signal processing. However current INR techniques suffer from a restricted capability to tune their supported frequency set resulting in i…

Cited by 34SourcePDFScholar
2024

NeRI: Implicit Neural Representation of LiDAR Point Cloud Using Range Image Sequence

ICASSP 2024accepted

This paper proposes the NeRI, an implicit neural representation (INR) based LiDAR point cloud compressor. In NeRI, we first transform a sequence of 3D LiDAR frames into a 2D range image sequence through range image projection over time. Then, we employ a neural network conditioned on the temporal fr…

Cited by 0SourceScholar
2024

PNeRV: Enhancing Spatial Consistency via Pyramidal Neural Representation for Videos

CVPR 2024poster

The primary focus of Neural Representation for Videos (NeRV) is to effectively model its spatiotemporal consistency. However current NeRV systems often face a significant issue of spatial inconsistency leading to decreased perceptual quality. To address this issue we introduce the Pyramidal Neural R…

Cited by 2SourcePDFScholar
2024

Towards Backward-Compatible Continual Learning of Image Compression

CVPR 2024poster

This paper explores the possibility of extending the capability of pre-trained neural image compressors (e.g. adapting to new data or target bitrates) without breaking backward compatibility the ability to decode bitstreams encoded by the original model. We refer to this problem as continual learnin…

2023

DINER: Disorder-Invariant Implicit Neural Representation

CVPR 2023highlight

Implicit neural representation (INR) characterizes the attributes of a signal as a function of corresponding coordinates which emerges as a sharp weapon for solving inverse problems. However, the capacity of INR is limited by the spectral bias in the network training. In this paper, we find that suc…

2023

DNeRV: Modeling Inherent Dynamics via Difference Neural Representation for Videos

CVPR 2023poster

Existing implicit neural representation (INR) methods do not fully exploit spatiotemporal redundancies in videos. Index-based INRs ignore the content-specific spatial features and hybrid INRs ignore the contextual dependency on adjacent frames, leading to poor modeling capability for scenes with lar…

Cited by 40SourcePDFScholar
2019

Spectral Reconstruction From Dispersive Blur: A Novel Light Efficient Spectral Imager

CVPR 2019poster

Developing high light efficiency imaging techniques to retrieve high dimensional optical signal is a long-term goal in computational photography. Multispectral imaging, which captures images of different wavelengths and boosting the abilities for revealing scene properties, has developed rapidly in…

Cited by 5PDFScholar
2016

Fast intra mode decision and block matching for HEVC screen content compression

ICASSP 2016accepted

Screen content coding (SCC) is the latest extension of the High-Efficiency Video Coding (HEVC) aiming to improve the compression efficiency of screen content video. With newly developed tools such as intra block copy (IntraBC) and palette (PLT) mode, SCC has been able to compress the desktop screens…

Cited by 0SourceScholar