← Search

Wen-Hsiao Peng

14 accepted papers

2025

Bridging Compressed Image Latents and Multimodal Large Language Models

ICLR 2025poster

This paper presents the first-ever study of adapting compressed image latents to suit the needs of downstream vision tasks that adopt Multimodal Large Language Models (MLLMs). MLLMs have extended the success of large language models to modalities (e.g. images) beyond text, but their billion scale hi…

Cited by 1SourcePDFScholar
2025

CAT-3DGS: A Context-Adaptive Triplane Approach to Rate-Distortion-Optimized 3DGS Compression

ICLR 2025poster

3D Gaussian Splatting (3DGS) has recently emerged as a promising 3D representation. Much research has been focused on reducing its storage requirements and memory footprint. However, the needs to compress and transmit the 3DGS representation to the remote side are overlooked. This new application ca…

Cited by 3SourcePDFScholar
2025

HyTIP: Hybrid Temporal Information Propagation for Masked Conditional Residual Video Coding

ICCV 2025poster

Most frame-based learned video codecs can be interpreted as recurrent neural networks (RNNs) propagating reference information along the temporal dimension. This work revisits the limitations of the current approaches from an RNN perspective. The output-recurrence methods, which propagate decoded fr…

2025

MH-LVC: Multi-Hypothesis Temporal Prediction for Learned Conditional Residual Video Coding

ICCV 2025poster

This work, termed MH-LVC, presents a multi-hypothesis temporal prediction scheme that employs long- and short-term reference frames in a conditional residual video coding framework. Recent temporal context mining approaches to conditional video coding offer superior coding performance. However, the…

2023

Hierarchical B-Frame Video Coding Using Two-Layer CANF Without Motion Coding

CVPR 2023highlight

Typical video compression systems consist of two main modules: motion coding and residual coding. This general architecture is adopted by classical coding schemes (such as international standards H.265 and H.266) and deep learning-based coding schemes. We propose a novel B-frame coding architecture…

2023

Learning Continuous Exposure Value Representations for Single-Image HDR Reconstruction

ICCV 2023poster

Deep learning is commonly used to produce impressive results in reconstructing HDR images from LDR images. LDR stack-based methods are used for single-image HDR reconstruction, generating an HDR image from a deep learning generated LDR stack. However, current methods generate the LDR stack with pred…

Cited by 10PDFScholar
2023

MoTIF: Learning Motion Trajectories with Local Implicit Neural Functions for Continuous Space-Time Video Super-Resolution

ICCV 2023poster

This work addresses continuous space-time video super-resolution (C-STVSR) that aims to up-scale an input video both spatially and temporally by any scaling factors. One key challenge of C-STVSR is to propagate information temporally among the input video frames. To this end, we introduce a space-ti…

Cited by 17PDFcodeScholar
2023

TransTIC: Transferring Transformer-based Image Compression from Human Perception to Machine Perception

ICCV 2023poster

This work aims for transferring a Transformer-based image compression codec from human perception to machine perception without fine-tuning the codec. We propose a transferable Transformer-based image compression framework, termed TransTIC. Inspired by visual prompt tuning, TransTIC adopts an instan…

Cited by 33PDFcodeScholar
2022

CANF-VC: Conditional Augmented Normalizing Flows for Video Compression

ECCV 2022poster

"This paper presents an end-to-end learning-based video compression system, termed CANF-VC, based on conditional augmented normalizing flows (CANF). Most learned video compression systems adopt the same hybrid-based coding architecture as the traditional codecs. Recent research on conditional coding…

2021

Video Rescaling Networks With Joint Optimization Strategies for Downscaling and Upscaling

CVPR 2021poster

This paper addresses the video rescaling task, which arises from the needs of adapting the video spatial resolution to suit individual viewing devices. We aim to jointly optimize video downscaling and upscaling as a combined task. Most recent studies focus on image-based solutions, which do not cons…

Cited by 17PDFcodeScholar
2019

All About Structure: Adapting Structural Information Across Domains for Boosting Semantic Segmentation

CVPR 2019poster

In this paper we tackle the problem of unsupervised domain adaptation for the task of semantic segmentation, where we attempt to transfer the knowledge learned upon synthetic datasets with ground-truth labels to real-world images without any annotation. With the hypothesis that the structural conten…

Cited by 317PDFcodeScholar
2019

SME-Net: Sparse Motion Estimation for Parametric Video Prediction Through Reinforcement Learning

ICCV 2019poster

This paper leverages a classic prediction technique, known as parametric overlapped block motion compensation (POBMC), in a reinforcement learning framework for video prediction. Learning-based prediction methods with explicit motion models often suffer from having to estimate large numbers of motio…

Cited by 16PDFScholar