← Search

Yi-Hsin Chen

9 accepted papers

2025

Bridging Compressed Image Latents and Multimodal Large Language Models

ICLR 2025poster

This paper presents the first-ever study of adapting compressed image latents to suit the needs of downstream vision tasks that adopt Multimodal Large Language Models (MLLMs). MLLMs have extended the success of large language models to modalities (e.g. images) beyond text, but their billion scale hi…

Cited by 1SourcePDFScholar
2025

CAT-3DGS: A Context-Adaptive Triplane Approach to Rate-Distortion-Optimized 3DGS Compression

ICLR 2025poster

3D Gaussian Splatting (3DGS) has recently emerged as a promising 3D representation. Much research has been focused on reducing its storage requirements and memory footprint. However, the needs to compress and transmit the 3DGS representation to the remote side are overlooked. This new application ca…

Cited by 3SourcePDFScholar
2025

HyTIP: Hybrid Temporal Information Propagation for Masked Conditional Residual Video Coding

ICCV 2025poster

Most frame-based learned video codecs can be interpreted as recurrent neural networks (RNNs) propagating reference information along the temporal dimension. This work revisits the limitations of the current approaches from an RNN perspective. The output-recurrence methods, which propagate decoded fr…

2025

MH-LVC: Multi-Hypothesis Temporal Prediction for Learned Conditional Residual Video Coding

ICCV 2025poster

This work, termed MH-LVC, presents a multi-hypothesis temporal prediction scheme that employs long- and short-term reference frames in a conditional residual video coding framework. Recent temporal context mining approaches to conditional video coding offer superior coding performance. However, the…

2023

MoTIF: Learning Motion Trajectories with Local Implicit Neural Functions for Continuous Space-Time Video Super-Resolution

ICCV 2023poster

This work addresses continuous space-time video super-resolution (C-STVSR) that aims to up-scale an input video both spatially and temporally by any scaling factors. One key challenge of C-STVSR is to propagate information temporally among the input video frames. To this end, we introduce a space-ti…

Cited by 17PDFcodeScholar
2023

TransTIC: Transferring Transformer-based Image Compression from Human Perception to Machine Perception

ICCV 2023poster

This work aims for transferring a Transformer-based image compression codec from human perception to machine perception without fine-tuning the codec. We propose a transferable Transformer-based image compression framework, termed TransTIC. Inspired by visual prompt tuning, TransTIC adopts an instan…

Cited by 33PDFcodeScholar
2022

ConTextING: Granting Document-Wise Contextual Embeddings to Graph Neural Networks for Inductive Text Classification

COLING 2022main

Graph neural networks (GNNs) have been recently applied in natural language processing. Various GNN research studies are proposed to learn node interactions within the local graph of each document that contains words, sentences, or topics for inductive text classification. However, most inductive GN…

2021

Video Rescaling Networks With Joint Optimization Strategies for Downscaling and Upscaling

CVPR 2021poster

This paper addresses the video rescaling task, which arises from the needs of adapting the video spatial resolution to suit individual viewing devices. We aim to jointly optimize video downscaling and upscaling as a combined task. Most recent studies focus on image-based solutions, which do not cons…

Cited by 17PDFcodeScholar
2017

No More Discrimination: Cross City Adaptation of Road Scene Segmenters

ICCV 2017poster

Despite the recent success of deep-learning based semantic segmentation, deploying a pre-trained road scene segmenter to a city whose images are not presented in the training set would not achieve satisfactory performance due to dataset biases. Instead of collecting a large number of annotated image…

Cited by 410PDFScholar