← Search

Zhiyong Gao

12 accepted papers

2023

Non-Semantics Suppressed Mask Learning for Unsupervised Video Semantic Compression

ICCV 2023poster

Most video compression methods aim to improve the decoded video visual quality, instead of particularly guaranteeing the semantic-completeness, which deteriorates downstream video analysis tasks, e.g., action recognition. In this paper, we focus on a novel unsupervised video semantic compression pro…

Cited by 27PDFcodeScholar
2021

Bounds all around: training energy-based models with bidirectional bounds

NeurIPS 2021poster

Energy-based models (EBMs) provide an elegant framework for density estimation, but they are notoriously difficult to train. Recent work has established links to generative adversarial networks, where the EBM is trained through a minimax game with a variational value function. We propose a bidirecti…

Cited by 19SourcePDFScholar
2021

Self-Conditioned Probabilistic Learning of Video Rescaling

ICCV 2021poster

Bicubic downscaling is a prevalent technique used to reduce the video storage burden or to accelerate the downstream processing speed. However, the inverse upscaling step is non-trivial, and the downscaled video may also deteriorate the performance of downstream tasks. In this paper, we propose a se…

Cited by 22PDFcodeScholar
2020

Content Adaptive and Error Propagation Aware Deep Video Compression

ECCV 2020poster

Recently, learning based video compression methods attract increasing attention. However, previous works suffer from error propagation, which stems from the accumulation of reconstructed error in inter predictive coding. Meanwhile, previous learning based video codecs are also not adaptive to differ…

Cited by 160SourcePDFScholar
2020

Self-supervised Motion Representation via Scattering Local Motion Cues

ECCV 2020poster

Motion representation is key to many computer vision problems but has never been well studied in the literature. Existing works usually rely on the optical flow estimation to assist other tasks such as action recognition, frame prediction, video segmentation, etc. In this paper, we leverage the mass…

Cited by 22SourcePDFScholar
2019

DVC: An End-To-End Deep Video Compression Framework

CVPR 2019oral

Conventional video compression approaches use the predictive coding architecture and encode the corresponding motion information and residual information. In this paper, taking advantage of both classical architecture in the conventional video compression method and the powerful non-linear represent…

Cited by 832PDFcodeScholar
2019

Depth-Aware Video Frame Interpolation

CVPR 2019poster

Video frame interpolation aims to synthesize nonexistent frames in-between the original frames. While significant advances have been made from the recent deep convolutional neural networks, the quality of interpolation is often reduced due to large object motion or occlusion. In this work, we propos…

Cited by 672PDFcodeScholar
2018

Deep Kalman Filtering Network for Video Compression Artifact Reduction

ECCV 2018poster

When lossy video compression algorithms are applied, compression artifacts often appear in videos, making decoded videos unpleasant for human visual systems. In this paper, we model the video artifact reduction task as a Kalman filtering procedure and restore decoded frames through a deep Kalman fil…

Cited by 120SourcePDFScholar
2018

Rcdfnn: Robust Change Detection Based on Convolutional Fusion Neural Network

ICASSP 2018accepted

Video change detection, which plays an important role in computer vision, is far from being well resolved due to the complexity of diverse scenes in real world. Most of the current methods are designed based on hand-crafted features and perform well in some certain scenes but may fail on others. Thi…

Cited by 0SourceScholar
2016

Principal components analysis-based visual saliency detection

ICASSP 2016accepted

In this paper, a novel patch-wise saliency detection algorithm is proposed based on Principal Component Analysis (PCA). As a powerful statistical procedure in data analysis, PCA are fully exploited to convert color space and produce compact patch representation. Specifically, images are first conver…

Cited by 0SourceScholar