← Search

Xinfeng Zhang

21 accepted papers

2026

GauMVC: Generative Decoupled Gaussian Representation for Human-centric Multi-view Video Compression

CVPR 2026

Human-centric multi-view video has a clear semantic structure: a static background and dynamic human motion. We propose a generative compression framework that explicitly decouples these components. The background is modeled once with 3D Gaussian Splatting, while the human is represented by a person

Cited by 0SourceScholar
2026

Spatio-Temporal Distortion Aware Omnidirectional Video Super-Resolution

AAAI 2026technical

Omnidirectional videos (ODVs) provide an immersive visual experience by capturing the 360° scene. With the rapid advancements in virtual/augmented reality, metaverse, and generative artificial intelligence, the demand for high-quality ODVs is surging. However, ODVs often suffer from low resolution d

Cited by 0SourcePDFScholar
2026

Spike Stream Memory Transfer for Dynamic Scene Reconstruction

AAAI 2026technical

As a retina-inspired sensor with ultra-high temporal resolution, spike camera can continuously capture dynamic scenes with high-speed motion. It is a key task to restore clear images from spike streams. The quantization effects in spike readout bring degradation to the visual quality of restored ima

Cited by 0SourcePDFScholar
2025

DialogDraw: Image Generation and Editing System Based on Multi-Turn Dialogue

AAAI 2025technical

In recent years, diffusion modeling has shown great potential for image generation and editing. Beyond single-model approaches, various drawing workflows now exist to handle diverse drawing tasks. However, few solutions effectively identify user intentions through dialogue and progressively complete…

Cited by 0SourcePDFScholar
2025

High Dynamic Range Imaging with Time-Encoding Spike Camera

NeurIPS 2025poster

As a bio-inspired vision sensor, spike camera records light intensity by accumulating photons and firing a spike once a preset threshold is reached. For high-light regions, the accumulated photons may reach the threshold multiple times within a readout interval, while only one spike can be stored an…

Cited by 0SourceScholar
2024

Optical Flow for Spike Camera with Hierarchical Spatial-Temporal Spike Fusion

AAAI 2024technical

As an emerging neuromorphic camera with an asynchronous working mechanism, spike camera shows good potential for high-speed vision tasks. Each pixel in spike camera accumulates photons persistently and fires a spike whenever the accumulation exceeds a threshold. Such high-frequency fine-granularity…

2024

Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-Modal Structured Representations

AAAI 2024technical

Large-scale vision-language pre-training has achieved significant performance in multi-modal understanding and generation tasks. However, existing methods often perform poorly on image-text matching tasks that require structured representations, i.e., representations of objects, attributes, and rela…

2023

Diffusion-Based 3D Human Pose Estimation with Multi-Hypothesis Aggregation

ICCV 2023poster

In this paper, a novel Diffusion-based 3D Pose estimation (D3DP) method with Joint-wise reProjection-based Multi-hypothesis Aggregation (JPMA) is proposed for probabilistic 3D human pose estimation. On the one hand, D3DP generates multiple possible 3D pose hypotheses for a single 2D observation. It…

Cited by 125PDFcodeScholar
2023

Recurrent Fine-Grained Self-Attention Network for Video Crowd Counting

ICASSP 2023accepted

Striking a balance between exploring the spatio-temporal correlation and controlling model complexity is vital for video-based crowd counting methods. In this paper, we propose a Recurrent Fine-Grained Self-Attention Network (RFSNet) to achieve efficient and accurate counting in video scenes via the…

Cited by 0SourceScholar
2022

P-STMO: Pre-trained Spatial Temporal Many-to-One Model for 3D Human Pose Estimation

ECCV 2022poster

"This paper introduces a novel Pre-trained Spatial Temporal Many-to-One (P-STMO) model for 2D-to-3D human pose estimation task. To reduce the difficulty of capturing spatial and temporal information, we divide this task into two stages: pre-training (Stage I) and fine-tuning (Stage II). In Stage I,…

2022

STRPM: A Spatiotemporal Residual Predictive Model for High-Resolution Video Prediction

CVPR 2022poster

Although many video prediction methods have obtained good performance in low-resolution (64 128) videos, predictive models for high-resolution (512 4K) videos have not been fully explored yet, which are more meaningful due to the increasing demand for high-quality videos. Compared with low-resolutio…

Cited by 68PDFScholar
2021

Evolutionary Quantization of Neural Networks with Mixed-Precision

ICASSP 2021accepted

Quantization is an effective way for reducing the memory and computation costs of deep neural networks. Most of existing methods exploit the fixed-precision quantization approach, e.g., weights and activations (i.e., output features) are represented as 8-bit values. Although mixed-precision quantiza…

Cited by 0SourceScholar
2021

MAU: A Motion-Aware Unit for Video Prediction and Beyond

NeurIPS 2021poster

Accurately predicting inter-frame motion information plays a key role in video prediction tasks. In this paper, we propose a Motion-Aware Unit (MAU) to capture reliable inter-frame motion information by broadening the temporal receptive field of the predictive units. The MAU consists of two modules,…

2021

Teacher-Student Learning With Multi-Granularity Constraint Towards Compact Facial Feature Representation

ICASSP 2021accepted

In this paper, we propose a novel end-to-end feature compression scheme by leveraging the representation and learning capability of deep neural networks, towards intelligent front-end equipped analysis with promising accuracy and efficiency. In particular, the extracted features are compactly coded…

Cited by 0SourceScholar
2020

Just Noticeable Distortion Based Perceptually Lossless Intra Coding

ICASSP 2020accepted

Perceptual video coding plays a very important role in video codec optimization aiming at removing the perceptual redundancies in video content. In this paper, a just noticeable distortion (JND) guided perceptually lossless coding framework is proposed for Versatile Video Coding (VVC) intra coding.…

Cited by 0SourceScholar
2019

A Data-centric Approach to Unsupervised Texture Segmentation Using Principle Representative Patterns

ICASSP 2019accepted

Features that capture textural patterns of a certain class of images are crucial for texture segmentation tasks. This paper introduces a data-centric approach to efficiently extract and represent textural information, which adapts to a wide variety of textures. Based on the strong self-similarities…

Cited by 0SourceScholar
2019

Cascaded Parallel Filtering for Memory-Efficient Image-Based Localization

ICCV 2019poster

Image-based localization (IBL) aims to estimate the 6DOF camera pose for a given query image. The camera pose can be computed from 2D-3D matches between a query image and Structure-from-Motion (SfM) models. Despite recent advances in IBL, it remains difficult to simultaneously resolve the memory con…

Cited by 38PDFcodeScholar
2018

Cluster-Based Point Cloud Coding with Normal Weighted Graph Fourier Transform

ICASSP 2018accepted

Point cloud has attracted more and more attention in 3D object representation, especially in free-view rendering. However, it is challenging to efficiently deploy the point cloud due to its huge data amount with multiple attributes including coordinates, normal and color. In order to represent point…

Cited by 0SourceScholar
2016

Aspect Ratio Similarity (ARS) for image retargeting quality assessment

ICASSP 2016accepted

During the past few years, there have been various kinds of content-aware image retargeting methods proposed for image resizing. However, the lack of effective objective retargeting quality metric limits the further development of image retargeting. Different from the traditional image quality asses…

Cited by 0SourceScholar