← Search

Chia-Wen Lin

32 accepted papers

2026

Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order Transfer

AAAI 2026technical

Action recognition using uncrewed aerial vehicles (UAVs) faces unique challenges due to substantial view variations along the vertical spatial axis. Unlike ground-based scenarios, UAVs capture actions from diverse altitudes, resulting in pronounced appearance discrepancies and reduced recognition ro

Cited by 0SourcePDFScholar
2025

4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding

ICCV 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated impressive 2D image/video understanding capabilities.However, there are no publicly standardized benchmarks to assess the abilities of MLLMs in understanding the 4D objects.In this paper, we introduce 4D-Bench, the first benchmark to evaluat…

2025

BlurDM: A Blur Diffusion Model for Image Deblurring

NeurIPS 2025poster

Diffusion models show promise for dynamic scene deblurring; however, existing studies often fail to leverage the intrinsic nature of the blurring process within diffusion models, limiting their full potential. To address it, we present a Blur Diffusion Model (BlurDM), which seamlessly integrates the…

Cited by 0SourcecodeScholar
2025

Generation and Comprehension Hand-in-Hand: Vision-guided Expression Diffusion for Boosting Referring Expression Generation and Comprehension

ICLR 2025poster

Referring expression generation (REG) and comprehension (REC) are vital and complementary in joint visual and textual reasoning. Existing REC datasets typically contain insufficient image-expression pairs for training, hindering the generalization of REC models to unseen referring expressions. More…

Cited by 0SourcePDFScholar
2025

PHATNet: A Physics-guided Haze Transfer Network for Domain-adaptive Real-world Image Dehazing

ICCV 2025poster

Image dehazing aims to remove unwanted hazy artifacts in images. Although previous research has collected paired real-world hazy and haze-free images to improve dehazing models' performance in real-world scenarios, these models often experience significant performance drops when handling unseen real…

2025

Spiking Meets Attention: Efficient Remote Sensing Image Super-Resolution with Attention Spiking Neural Networks

NeurIPS 2025poster

Spiking neural networks (SNNs) are emerging as a promising alternative to traditional artificial neural networks (ANNs), offering biological plausibility and energy efficiency. Despite these merits, SNNs are frequently hampered by limited capacity and insufficient representation power, yet remain un…

Cited by 0SourcecodeScholar
2025

Towards General Visual-Linguistic Face Forgery Detection

CVPR 2025poster

Face manipulation techniques have achieved significant advances, presenting serious challenges to security and social trust. Recent works demonstrate that leveraging multimodal models can enhance the generalization and interpretability of face forgery detection. However, existing annotation approach…

2024

A Fine-Grained Attribute Pre-Labeling Method Based on Label Dependency and Feature Similarity Dynamics

ICASSP 2024accepted

In this paper, we proposed a fine-grained attribute pre-labeling method based on the multi-label recovery techniques. Given a fine-grained image dataset with overlooked attributes in its annotation vectors, our method can predict those missing attribute labels by learning the between-label dependenc…

Cited by 0SourceScholar
2024

Domain-adaptive Video Deblurring via Test-time Blurring

ECCV 2024poster

"Dynamic scene video deblurring aims to remove undesirable blurry artifacts captured during the exposure process. Although previous video deblurring methods have achieved impressive results, they suffer from significant performance drops due to the domain gap between training and testing videos, esp…

2024

ID-Blau: Image Deblurring by Implicit Diffusion-based reBLurring AUgmentation

CVPR 2024poster

Image deblurring aims to remove undesired blurs from an image captured in a dynamic scene. Much research has been dedicated to improving deblurring performance through model architectural designs. However there is little work on data augmentation for image deblurring. Since continuous motion causes…

2023

NewsNet: A Novel Dataset for Hierarchical Temporal Segmentation

CVPR 2023poster

Temporal video segmentation is the get-to-go automatic video analysis, which decomposes a long-form video into smaller components for the following-up understanding tasks. Recent works have studied several levels of granularity to segment a video, such as shot, event, and scene. Those segmentations…

2023

Only a Few Classes Confusing: Pixel-Wise Candidate Labels Disambiguation for Foggy Scene Understanding

AAAI 2023technical

Not all semantics become confusing when deploying a semantic segmentation model for real-world scene understanding of adverse weather. The true semantics of most pixels have a high likelihood of appearing in the few top classes according to confidence ranking. In this paper, we replace the one-hot p…

Cited by 9SourcePDFScholar
2022

Both Style and Fog Matter: Cumulative Domain Adaptation for Semantic Foggy Scene Understanding

CVPR 2022oral

Although considerable progress has been made in semantic scene understanding under clear weather, it is still a tough problem under adverse weather conditions, such as dense fog, due to the uncertainty caused by imperfect observations. Besides, difficulties in collecting and labeling foggy images hi…

Cited by 64PDFScholar
2022

DANet: Image Deraining via Dynamic Association Learning

IJCAI 2022poster

Rain streaks and background components in a rainy input are highly correlated, making the deraining task a composition of the rain streak removal and background restoration. However, the correlation of these two components is barely considered, leading to unsatisfied deraining results. To this end,…

Cited by 21SourcePDFScholar
2022

Degrade Is Upgrade: Learning Degradation for Low-Light Image Enhancement

AAAI 2022technical

Low-light image enhancement aims to improve an image's visibility while keeping its visual naturalness. Different from existing methods, which tend to accomplish the relighting task directly, we investigate the intrinsic degradation and relight the low-light image while refining the details and colo…

2022

Fast Graph Sampling for Short Video Summarization Using Gershgorin Disc Alignment

ICASSP 2022accepted

We study the problem of efficiently summarizing a short video into several keyframes, leveraging recent progress in fast graph sampling. Specifically, we first construct a similarity path graph (SPG) G, represented by graph Laplacian matrix L, where the similarities between adjacent frames are encod…

Cited by 0SourceScholar
2022

Seeing through a Black Box: Toward High-Quality Terahertz Imaging via Subspace-and-Attention Guided Restoration

ECCV 2022poster

"Terahertz (THz) imaging has recently attracted significant attention thanks to its non-invasive, non-destructive, non-ionizing, material-classification, and ultra-fast nature for object exploration and inspection. However, its strong water absorption nature and low noise tolerance lead to undesired…

Cited by 3SourcePDFScholar
2022

Stripformer: Strip Transformer for Fast Image Deblurring

ECCV 2022poster

"Images taken in dynamic scenes may contain unwanted motion blur, which significantly degrades visual quality. Such blur causes short- and long-range region-specific smoothing artifacts that are often directional and non-uniform, which is difficult to be removed. Inspired by the current success of t…

2021

Discover Cross-Modality Nuances for Visible-Infrared Person Re-Identification

CVPR 2021poster

Visible-infrared person re-identification (Re-ID) aims to match the pedestrian images of the same identity from different modalities. Existing works mainly focus on alleviating the modality discrepancy by aligning the distributions of features from different modalities. However, nuanced but discrimi…

Cited by 286PDFcodeScholar
2021

Dual-level Collaborative Transformer for Image Captioning

AAAI 2021technical

Descriptive region features extracted by object detection networks have played an important role in the recent advancements of image captioning. However, they are still criticized for the lack of contextual information and fine-grained details, which in contrast are the merits of traditional grid fe…

2021

High Quality Disparity Remapping With Two-Stage Warping

ICCV 2021poster

A high quality disparity remapping method that preserves 2D shapes and 3D structures, and adjusts disparities of important objects in stereo image pairs is proposed. It is formulated as a constrained optimization problem, whose solution is challenging, since we need to meet multiple requirements of…

Cited by 2PDFScholar
2021

Image Inpainting Guided by Coherence Priors of Semantics and Textures

CVPR 2021poster

Existing inpainting methods have achieved promising performance in recovering defected images of specific scenes. However, filling holes involving multiple semantic categories remains challenging due to the obscure semantic boundaries and the mixture of different semantic textures. In this paper, we…

Cited by 120PDFScholar
2020

Graph Neural Net Using Analytical Graph Filters and Topology Optimization for Image Denoising

ICASSP 2020accepted

While convolutional neural nets (CNNs) have achieved remarkable performance for a wide range of inverse imaging applications, the filter coefficients are computed in a purely data-driven manner and are not explainable. Inspired by an analytically derived CNN by Hadji et al., in this paper we constru…

Cited by 0SourceScholar
2020

Guidance and Evaluation: Semantic-Aware Image Inpainting for Mixed Scenes

ECCV 2020poster

Completing a corrupted image with correct structures and reasonable textures for a mixed scene remains an elusive challenge. Since the missing hole in a mixed scene of a corrupted image often contains various semantic information, conventional two-stage approaches utilizing structural information of…

Cited by 154SourcePDFScholar
2020

HardGAN: A Haze-Aware Representation Distillation GAN for Single Image Dehazing

ECCV 2020poster

In this paper, we present a Haze-Aware Representation Distillation Generative Adversarial Network named HardGAN for single-image dehazing. Unlike previous studies that intend to model the transmission map and global atmospheric light jointly to restore a clear image, we solve this regression problem…

2020

Rotated Binary Neural Network

NeurIPS 2020poster

Binary Neural Network (BNN) shows its predominance in reducing the complexity of deep neural networks. However, it suffers severe performance degradation. One of the major impediments is the large quantization error between the full-precision weight vector and its binary vector. Previous works focus…

2020

When Pedestrian Detection Meets Nighttime Surveillance: A New Benchmark

IJCAI 2020poster

Pedestrian detection at nighttime is a crucial and frontier problem in surveillance, but has not been well explored by the computer vision and artificial intelligence communities. Most of existing methods detect pedestrians under favorable lighting conditions (e.g. daytime) and achieve promising per…

2019

Consistency Constrained Reconstruction of Depth Maps from Epipolar Plane Images

ICASSP 2019accepted

In this paper, we propose a method of reconstructing the depth map of a set of multiview images from the epipolar plane images (EPIs) of multiview Images. Our method involves two steps: finding support points and estimating depth. First, we propose to include a consistency term and a smoothness term…

Cited by 0SourceScholar
2019

Information Competing Process for Learning Diversified Representations

NeurIPS 2019poster

Learning representations with diversified information remains as an open problem. Towards learning diversified representations, a new approach, termed Information Competing Process (ICP), is proposed in this paper. Aiming to enrich the information carried by feature representations, ICP separates a…

2016

Supervised-learning based face hallucination for enhancing face recognition

ICASSP 2016accepted

This paper presents a two-step supervised face hallucination framework based on class-specific dictionary learning. Since the performance of learning-based face hallucination relies on its training set, an inappropriate training set (e.g., an input face image is very different from the training set)…

Cited by 0SourceScholar