← Search

Wen Gao

34 accepted papers

2026

Discovering Adaptive Task Dependencies for Efficient Multi-Task Representation Compression

CVPR 2026

Traditional image compression prioritizes pixel fidelity but often preserves details irrelevant to downstream vision tasks. Compressing task-specific representations instead better aligns with task semantics, yet redundant information persists across correlated tasks. Existing multi-task compression

Cited by 0SourceScholar
2025

CALLIC: Content Adaptive Learning for Lossless Image Compression

AAAI 2025technical

Learned lossless image compression has achieved significant advancements in recent years. However, existing methods often rely on training amortized generative models on massive datasets, resulting in sub-optimal probability distribution estimation for specific testing images during encoding process…

Cited by 1SourcePDFScholar
2025

Emerging Advances in Learned Video Compression: Models, Systems and Beyond

IJCAI 2025

Video compression is a fundamental topic in the visual intelligence, bridging visual signal sensing/capturing and high-level visual analytics. The broad success of artificial intelligence (AI) technology has enriched the horizon of video compression into novel paradigms by leveraging end-to-end opti

Cited by 0SourcePDFScholar
2025

Hierarchical Deep Reinforcement Learning for Computation Offloading in Autonomous Multi-Robot Systems

RA-L 2025

To ensure system responsiveness, some compute-intensive tasks are usually offloaded to cloud or edge computing devices. In environments where connection to external computing facilities is unavailable, computation offloading among members within an autonomous multi-robot system (AMRS) becomes a solu

Cited by 4SourceScholar
2023

Diffusion-Based 3D Human Pose Estimation with Multi-Hypothesis Aggregation

ICCV 2023poster

In this paper, a novel Diffusion-based 3D Pose estimation (D3DP) method with Joint-wise reProjection-based Multi-hypothesis Aggregation (JPMA) is proposed for probabilistic 3D human pose estimation. On the one hand, D3DP generates multiple possible 3D pose hypotheses for a single 2D observation. It…

Cited by 125PDFcodeScholar
2022

P-STMO: Pre-trained Spatial Temporal Many-to-One Model for 3D Human Pose Estimation

ECCV 2022poster

"This paper introduces a novel Pre-trained Spatial Temporal Many-to-One (P-STMO) model for 2D-to-3D human pose estimation task. To reduce the difficulty of capturing spatial and temporal information, we divide this task into two stages: pre-training (Stage I) and fine-tuning (Stage II). In Stage I,…

2022

STRPM: A Spatiotemporal Residual Predictive Model for High-Resolution Video Prediction

CVPR 2022poster

Although many video prediction methods have obtained good performance in low-resolution (64 128) videos, predictive models for high-resolution (512 4K) videos have not been fully explored yet, which are more meaningful due to the increasing demand for high-quality videos. Compared with low-resolutio…

Cited by 68PDFScholar
2022

Towards End-to-End Image Compression and Analysis with Transformers

AAAI 2022technical

We propose an end-to-end image compression and analysis model with Transformers, targeting to the cloud-based image classification application. Instead of placing an existing Transformer-based image classification model directly after an image codec, we aim to redesign the Vision Transformer (ViT) m…

2021

An Adaptive Pyramid Single-View Depth Lookup Table Coding Method

ICASSP 2021accepted

As depth maps show unique characteristics like piecewise smooth regions bounded by sharp edges at depth discontinuities, new coding tools are required to approximate these signal characteristics. Moreover, the number of bits to signal the residual values for each segment can be further reduced by in…

Cited by 0SourceScholar
2021

Evolutionary Quantization of Neural Networks with Mixed-Precision

ICASSP 2021accepted

Quantization is an effective way for reducing the memory and computation costs of deep neural networks. Most of existing methods exploit the fixed-precision quantization approach, e.g., weights and activations (i.e., output features) are represented as 8-bit values. Although mixed-precision quantiza…

Cited by 0SourceScholar
2021

MAU: A Motion-Aware Unit for Video Prediction and Beyond

NeurIPS 2021poster

Accurately predicting inter-frame motion information plays a key role in video prediction tasks. In this paper, we propose a Motion-Aware Unit (MAU) to capture reliable inter-frame motion information by broadening the temporal receptive field of the predictive units. The MAU consists of two modules,…

2021

Post-Training Quantization for Vision Transformer

NeurIPS 2021poster

Recently, transformer has achieved remarkable performance on a variety of computer vision applications. Compared with mainstream convolutional neural networks, vision transformers are often of sophisticated architectures for extracting powerful feature representations, which are more difficult to be…

Cited by 416SourcePDFScholar
2021

Pre-Trained Image Processing Transformer

CVPR 2021poster

As the computing power of modern hardware is increasing strongly, pre-trained deep learning models (e.g., BERT, GPT-3) learned on large-scale datasets have shown their effectiveness over conventional methods. The big progress is mainly contributed to the representation ability of transformer and its…

Cited by 2279PDFcodeScholar
2021

Progressive Stage-Wise Learning for Unsupervised Feature Representation Enhancement

CVPR 2021poster

Unsupervised learning methods have recently shown their competitiveness against supervised training. Typically, these methods use a single objective to train the entire network. But one distinct advantage of unsupervised over supervised learning is that the former possesses more variety and freedom…

Cited by 6PDFScholar
2021

Segatron: Segment-Aware Transformer for Language Modeling and Understanding

AAAI 2021technical

Transformers are powerful for sequence modeling. Nearly all state-of-the-art language models and pre-trained language models are based on the Transformer architecture. However, it distinguishes sequential tokens only with the token position index. We hypothesize that better contextual representation…

2019

Global-Local Temporal Representations for Video Person Re-Identification

ICCV 2019poster

This paper proposes the Global-Local Temporal Representation (GLTR) to exploit the multi-scale temporal cues in video sequences for video person Re-Identification (ReID). GLTR is constructed by first modeling the short-term temporal cues among adjacent frames, then capturing the long-term relations…

Cited by 286PDFScholar
2019

Multiscale Directional Fusion for Depth Map Super Resolution with Denoising

ICASSP 2019accepted

To tackle three main problems in depth map super resolution (SR) process, which are texture copy artifacts, blurred edge artifacts and jagged edge artifacts, we propose a depth map super resolution with denoising method based on multiscale directional fusion via nonsubsampled contourlet transform (N…

Cited by 0SourceScholar
2019

Reconstruction-cognizant Graph Sampling Using Gershgorin Disc Alignment

ICASSP 2019accepted

Graph sampling with noise is a fundamental problem in graph signal processing (GSP). Previous works assume an unbiased least square (LS) signal reconstruction scheme and select samples greedily via expensive extreme eigenvector computation. A popular biased scheme using graph Laplacian regularizatio…

Cited by 0SourceScholar
2018

Cluster-Based Point Cloud Coding with Normal Weighted Graph Fourier Transform

ICASSP 2018accepted

Point cloud has attracted more and more attention in 3D object representation, especially in free-view rendering. However, it is challenging to efficiently deploy the point cloud due to its huge data amount with multiple attributes including coordinates, normal and color. In order to represent point…

Cited by 0SourceScholar
2018

Person Transfer GAN to Bridge Domain Gap for Person Re-Identification

CVPR 2018poster

Although the performance of person Re-Identification (ReID) has been significantly boosted, many challenging issues in real scenarios have not been fully investigated, e.g., the complex scenes and lighting variations, viewpoint and pose changes, and the large number of identities in a camera network…

Cited by 2297SourcePDFScholar
2017

A cache-based bandwidth optimized motion compensation architecture for video decoder

ICASSP 2017accepted

In video decoder applications, motion compensation (MC) is bandwidth consuming because of the non-regular memory access. Especially with the popularity of UHD video and the development of new coding standard (HEVC), external memory bandwidth becomes a crucial bottleneck. In this paper, we propose an…

Cited by 0SourceScholar
2017

Performance Guaranteed Network Acceleration via High-Order Residual Quantization

ICCV 2017poster

Input binarization has shown to be an effective way for network acceleration. However, previous binarization scheme could be regarded as simple pixel-wise thresholding operations (i.e., order-one approximation) and suffers a big accuracy loss. In this paper, we propose a high-order binarization sche…

Cited by 137PDFScholar
2017

Pose-Driven Deep Convolutional Model for Person Re-Identification

ICCV 2017poster

Feature extraction and matching are two crucial components in person Re-Identification (ReID). The large pose deformations and the complex view variations exhibited by the captured person images significantly increase the difficulty of learning and matching of the features from person images. To ove…

Cited by 999PDFScholar
2016

AnalogCast: Full linear coding and pseudo analog transmission for satellite remote-sensing images

ICASSP 2016accepted

In this paper, we propose a novel image coding and transmission scheme called AnalogCast, which is a pseudo analog coding system for transmitting satellite remote-sensing images to large number of receivers. AnalogCast follows the idea originally developed for SoftCast [1-3] but with two special tec…

Cited by 0SourceScholar
2016

Deep Alternative Neural Network: Exploring Contexts as Early as Possible for Action Recognition

NeurIPS 2016poster

Contexts are crucial for action recognition in video. Current methods often mine contexts after extracting hierarchical local features and focus on their high-order encodings. This paper instead explores contexts as early as possible and leverages their evolutions for action recognition. In particul…

Cited by 27SourcePDFScholar
2015

Image Denoising via Adaptive Soft-Thresholding Based on Non-Local Samples

CVPR 2015poster

This paper proposes a new image denoising approach using adaptive signal modeling and adaptive soft-thresholding. It improves the image quality by regularizing all the patches in image based on distribution modeling in transform domain. Instead of using a global model for all patches, it employs con…

Cited by 97SourcePDFScholar
2015

Multi-Task Learning With Low Rank Attribute Embedding for Person Re-Identification

ICCV 2015poster

We propose a novel Multi-Task Learning with Low Rank Attribute Embedding (MTL-LORAE) framework for person re-identification. Re-identifications from multiple cameras are regarded as related tasks to exploit shared information to improve re-identification accuracy. Both low level features and semanti…

Cited by 204PDFScholar