← Search

Zhening Liu

11 accepted papers

2026

Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation

ICML 2026poster

Latent diffusion models have enabled high-quality video synthesis, yet their inference remains costly and time-consuming. As diffusion transformers become increasingly efficient, the latency bottleneck inevitably shifts to VAE decoders. To reduce their latency while maintaining quality, we propose a…

Cited by 0SourceScholar
2026

Low-Latency Neural LiDAR Compression with 2D Context Models

ICLR 2026poster

Context modeling is fundamental to LiDAR point cloud compression. Existing methods rely on computationally intensive 3D contexts, such as voxel and octree, which struggle to balance the compression efficiency and coding speed. In this work, we propose a neural LiDAR compressor based on 2D context mo…

Cited by 0SourcecodeScholar
2026

MambaSIC: Mamba-based Stereo Image Compression with Bi-directional Multi-reference Entropy Model

CVPR 2026

Stereo image compression (SIC) has become increasingly vital with its applications surging in fields such as 3D reconstruction and autonomous navigation. Previous methods leverage cross-attention to model inter-view redundancy and employ autoregressive entropy models to predict probability distribut

Cited by 0SourceScholar
2026

RemedyGS: Defend 3D Gaussian Splatting Against Computation Cost Attacks

CVPR 2026

As a mainstream technique for 3D reconstruction, 3D Gaussian splatting (3DGS) has been applied in a wide range of applications and services. Recent studies have revealed critical vulnerabilities in this pipeline and introduced computation cost attacks that lead to malicious resource occupancies and

Cited by 0SourcecodeScholar
2026

Spatia: Video Generation with Updatable Spatial Memory

CVPR 2026

Existing video generation models struggle to maintain long-term spatial and temporal consistency due to the dense, high-dimensional nature of video signals. To overcome this limitation, we propose Spatia, a spatial memory-aware video generation framework that explicitly preserves a 3D scene point cl

Cited by 0SourcecodeScholar
2025

CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression

AAAI 2025technical

Existing learning-based stereo image codec adopt sophisticated transformation with simple entropy models derived from single image codecs to encode latent representations. However, those entropy models struggle to effectively capture the spatial-disparity characteristics inherent in stereo images, w…

2025

MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic Scenes

ICCV 2025poster

4D Gaussian Splatting (4DGS) has recently emerged as a promising technique for capturing complex dynamic 3D scenes with high fidelity. It utilizes a 4D Gaussian representation and a GPU-friendly rasterizer, enabling rapid rendering speeds. Despite its advantages, 4DGS faces significant challenges, n…

2025

PF-TEB: Timed Elastic Band-Based Human-Aware Robot Navigation Framework in Crowded Environments

RA-L 2025

To enhance the social navigation performance of mobile robots in crowded environments, we propose a novel framework—Prediction and Fuzzy Timed Elastic Band (PF-TEB) for robot social navigation. Our framework incorporates predicted pedestrian trajectories into pedestrian proxemics modeling as a socia

Cited by 2SourceScholar
2024

Bidirectional Stereo Image Compression with Cross-Dimensional Entropy Model

ECCV 2024poster

"With the rapid advancement of stereo vision technologies, stereo image compression has emerged as a crucial field that continues to draw significant attention. Previous approaches have primarily employed a unidirectional paradigm, where the compression of one view is dependent on the other, resulti…

2024

STAGP: Spatio-Temporal Adaptive Graph Pooling Network for Pedestrian Trajectory Prediction

RA-L 2024

Predicting how pedestrians will move in the future is crucial for robot navigation, autonomous driving, and video surveillance. The complex interactions among pedestrians make it difficult to predict their future trajectory. Previous studies have primarily focused on modeling the interaction feature

Cited by 23SourceScholar