← Search

Yen-Yu Lin

39 accepted papers

2026

GOT-Edit: Geometry-Aware Generic Object Tracking via Online Model Editing

ICLR 2026poster

Human perception for effective object tracking in a 2D video stream arises from the implicit use of prior 3D knowledge combined with semantic reasoning. In contrast, most generic object tracking (GOT) methods primarily rely on 2D features of the target and its surroundings while neglecting 3D geomet…

Cited by 0SourcecodeScholar
2025

AuraFusion360: Augmented Unseen Region Alignment for Reference-based 360deg Unbounded Scene Inpainting

CVPR 2025poster

Three-dimensional scene inpainting is crucial for applications from virtual reality to architectural visualization, yet existing methods struggle with view consistency and geometric accuracy in 360deg unbounded scenes. We present AuraFusion360, a novel reference-based method that enables high-qualit…

Cited by 0SourcePDFScholar
2025

BlurDM: A Blur Diffusion Model for Image Deblurring

NeurIPS 2025poster

Diffusion models show promise for dynamic scene deblurring; however, existing studies often fail to leverage the intrinsic nature of the blurring process within diffusion models, limiting their full potential. To address it, we present a Blur Diffusion Model (BlurDM), which seamlessly integrates the…

Cited by 0SourcecodeScholar
2025

Generation and Comprehension Hand-in-Hand: Vision-guided Expression Diffusion for Boosting Referring Expression Generation and Comprehension

ICLR 2025poster

Referring expression generation (REG) and comprehension (REC) are vital and complementary in joint visual and textual reasoning. Existing REC datasets typically contain insufficient image-expression pairs for training, hindering the generalization of REC models to unseen referring expressions. More…

Cited by 0SourcePDFScholar
2025

LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos

ICCV 2025poster

LongSplat addresses critical challenges in novel view synthesis (NVS) from casually captured long videos characterized by irregular camera motion, unknown camera poses, and expansive scenes. Current methods often suffer from pose drift, inaccurate geometry initialization, and severe memory limitatio…

2025

PHATNet: A Physics-guided Haze Transfer Network for Domain-adaptive Real-world Image Dehazing

ICCV 2025poster

Image dehazing aims to remove unwanted hazy artifacts in images. Although previous research has collected paired real-world hazy and haze-free images to improve dehazing models' performance in real-world scenarios, these models often experience significant performance drops when handling unseen real…

2025

Ranking-aware adapter for text-driven image ordering with CLIP

ICLR 2025poster

Recent advances in vision-language models (VLMs) have made significant progress in downstream tasks that require quantitative concepts such as facial age estimation and image quality assessment, enabling VLMs to explore applications like image ranking and retrieval. However, existing studies typical…

2024

Domain-adaptive Video Deblurring via Test-time Blurring

ECCV 2024poster

"Dynamic scene video deblurring aims to remove undesirable blurry artifacts captured during the exposure process. Although previous video deblurring methods have achieved impressive results, they suffer from significant performance drops due to the domain gap between training and testing videos, esp…

2024

ID-Blau: Image Deblurring by Implicit Diffusion-based reBLurring AUgmentation

CVPR 2024poster

Image deblurring aims to remove undesired blurs from an image captured in a dynamic scene. Much research has been dedicated to improving deblurring performance through model architectural designs. However there is little work on data augmentation for image deblurring. Since continuous motion causes…

2024

Image-Text Co-Decomposition for Text-Supervised Semantic Segmentation

CVPR 2024poster

This paper addresses text-supervised semantic segmentation aiming to learn a model capable of segmenting arbitrary visual concepts within images by using only image-text pairs without dense annotations. Existing methods have demonstrated that contrastive learning on image-text pairs effectively alig…

2024

PartDistill: 3D Shape Part Segmentation by Vision-Language Model Distillation

CVPR 2024poster

This paper proposes a cross-modal distillation framework PartDistill which transfers 2D knowledge from vision-language models (VLMs) to facilitate 3D shape part segmentation. PartDistill addresses three major challenges in this task: the lack of 3D segmentation in invisible or undetected regions in…

2023

2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level Supervision

ICCV 2023poster

We present a Multimodal Interlaced Transformer (MIT) that jointly considers 2D and 3D data for weakly supervised point cloud segmentation. Research studies have shown that 2D and 3D features are complementary for point cloud segmentation. However, existing methods require extra 2D annotations to ach…

Cited by 16PDFScholar
2023

Diffusion-SS3D: Diffusion Model for Semi-supervised 3D Object Detection

NeurIPS 2023poster

Semi-supervised object detection is crucial for 3D scene understanding, efficiently addressing the limitation of acquiring large-scale 3D bounding box annotations. Existing methods typically employ a teacher-student framework with pseudo-labeling to leverage unlabeled point clouds. However, producin…

2023

Learning Continuous Exposure Value Representations for Single-Image HDR Reconstruction

ICCV 2023poster

Deep learning is commonly used to produce impressive results in reconstructing HDR images from LDR images. LDR stack-based methods are used for single-image HDR reconstruction, generating an HDR image from a deep learning generated LDR stack. However, current methods generate the LDR stack with pred…

Cited by 10PDFScholar
2023

MoTIF: Learning Motion Trajectories with Local Implicit Neural Functions for Continuous Space-Time Video Super-Resolution

ICCV 2023poster

This work addresses continuous space-time video super-resolution (C-STVSR) that aims to up-scale an input video both spatially and temporally by any scaling factors. One key challenge of C-STVSR is to propagate information temporally among the input video frames. To this end, we introduce a space-ti…

Cited by 17PDFcodeScholar
2022

AQT: Adversarial Query Transformers for Domain Adaptive Object Detection

IJCAI 2022poster

Adversarial feature alignment is widely used in domain adaptive object detection. Despite the effectiveness on CNN-based detectors, its applicability to transformer-based detectors is less studied. In this paper, we present AQT (adversarial query transformers) to integrate adversarial feature alignm…

2022

An MIL-Derived Transformer for Weakly Supervised Point Cloud Segmentation

CVPR 2022poster

We address weakly supervised point cloud segmentation by proposing a new model, MIL-derived transformer, to mine additional supervisory signals. First, the transformer model is derived based on multiple instance learning (MIL) to explore pair-wise cloud-level supervision, where two clouds of the sam…

Cited by 61PDFScholar
2022

Point MixSwap: Attentional Point Cloud Mixing via Swapping Matched Structural Divisions

ECCV 2022poster

"Data augmentation is developed for increasing the amount and diversity of training data to enhance model learning. Compared to 2D images, point clouds, with the 3D geometric nature as well as the high collection and annotation costs, pose great challenges and potentials for augmentation. This paper…

2022

Stripformer: Strip Transformer for Fast Image Deblurring

ECCV 2022poster

"Images taken in dynamic scenes may contain unwanted motion blur, which significantly degrades visual quality. Such blur causes short- and long-range region-specific smoothing artifacts that are often directional and non-uniform, which is difficult to be removed. Inspired by the current success of t…

2021

Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video Parsing

NeurIPS 2021poster

The audio-visual video parsing task aims to temporally parse a video into audio or visual event categories. However, it is labor intensive to temporally annotate audio and visual events and thus hampers the learning of a parsing model. To this end, we propose to explore additional cross-video and cr…

2021

Unsupervised Point Cloud Object Co-Segmentation by Co-Contrastive Learning and Mutual Attention Sampling

ICCV 2021poster

This paper presents a new task, point cloud object co-segmentation, aiming to segment the common 3D objects in a set of point clouds. We formulate this task as an object point sampling problem, and develop two techniques, the mutual attention module and co-contrastive learning, to enable it. The pro…

Cited by 17PDFcodeScholar
2020

Every Pixel Matters: Center-aware Feature Alignment for Domain Adaptive Object Detector

ECCV 2020poster

A domain adaptive object detector aims to adapt itself to unseen domains that may contain variations of object appearance, viewpoints or backgrounds. Most existing solutions adopt feature alignment either on the image level or instance level. However, image-level alignment on global features may tan…

2020

Spatiotemporal Super-Resolution with Cross-Task Consistency and Its Semi-supervised Extension

IJCAI 2020poster

Spatiotemporal super-resolution (SR) aims to upscale both the spatial and temporal dimensions of input videos, and produces videos with higher frame resolutions and rates. It involves two essential sub-tasks: spatial SR and temporal SR. We design a two-stream network for spatiotemporal SR in this wo…

2019

CrDoCo: Pixel-Level Domain Transfer With Cross-Domain Consistency

CVPR 2019poster

Unsupervised domain adaptation algorithms aim to transfer the knowledge learned from one domain to another (e.g., synthetic to real images). The adapted representations often do not capture pixel-level domain shifts that are crucial for dense prediction tasks (e.g., semantic segmentation). In this p…

Cited by 378PDFScholar
2019

DeepCO3: Deep Instance Co-Segmentation by Co-Peak Search and Co-Saliency Detection

CVPR 2019oral

In this paper, we address a new task called instance co-segmentation. Given a set of images jointly covering object instances of a specific category, instance co-segmentation aims to identify all of these instances and segment each of them, i.e. generating one mask for each instance. This task is im…

Cited by 86PDFcodeScholar
2019

FSA-Net: Learning Fine-Grained Structure Aggregation for Head Pose Estimation From a Single Image

CVPR 2019poster

This paper proposes a method for head pose estimation from a single image. Previous methods often predict head poses through landmark or depth estimation and would require more computation than necessary. Our method is based on regression and feature aggregation. For having a compact model, we emplo…

Cited by 384PDFcodeScholar
2019

Recover and Identify: A Generative Dual Model for Cross-Resolution Person Re-Identification

ICCV 2019poster

Person re-identification (re-ID) aims at matching images of the same identity across camera views. Due to varying distances between cameras and persons of interest, resolution mismatch can be expected, which would degrade person re-ID performance in real-world scenarios. To overcome this problem, we…

Cited by 97PDFScholar
2019

Weakly Supervised Instance Segmentation using the Bounding Box Tightness Prior

NeurIPS 2019poster

This paper presents a weakly supervised instance segmentation method that consumes training data with tight bounding box annotations. The major difficulty lies in the uncertain figure-ground separation within each bounding box since there is no supervisory signal about it. We address the difficulty…

2018

Unsupervised CNN-based Co-Saliency Detection with Graphical Optimization

ECCV 2018poster

In this paper, we address co-saliency detection in a set of images jointly covering objects of a specific class by an unsupervised convolutional neural network (CNN). Our method does not require any additional training data in the form of object masks. We decompose co-saliency detection into two sub…

Cited by 68SourcePDFScholar
2017

Deep Co-Occurrence Feature Learning for Visual Object Recognition

CVPR 2017poster

This paper addresses three issues in integrating part-based representations into convolutional neural networks (CNNs) for object recognition. First, most part-based models rely on a few pre-specified object parts. However, the optimal object parts for recognition often vary from category to category…

Cited by 53PDFcodeScholar
2017

DeepCD: Learning Deep Complementary Descriptors for Patch Representations

ICCV 2017poster

This paper presents the DeepCD framework which learns a pair of complementary descriptors jointly for a patch by employing deep learning techniques. It can be achieved by taking any descriptor learning architecture for learning a leading descriptor and augmenting the architecture with an additional…

Cited by 49PDFcodeScholar
2017

Learning and inferring human actions with temporal pyramid features based on conditional random fields

ICASSP 2017accepted

Finding an effective way to represent human actions is yet an open problem because it usually requires taking evidences extracted from various temporal resolutions into account. A conventional way of representing an action employs temporally ordered fine-grained movements, e.g., key poses or subtle…

Cited by 0SourceScholar
2016

Accumulated Stability Voting: A Robust Descriptor From Descriptors of Multiple Scales

CVPR 2016poster

This paper proposes a novel local descriptor through accumulated stability voting (ASV). The stability of feature dimensions is measured by their differences across scales. To be more robust to noise, the stability is further quantized by thresholding. The principle of maximum entropy is utilized fo…

Cited by 29PDFcodeScholar
2016

Precise player segmentation in team sports videos using contrast-aware co-segmentation

ICASSP 2016accepted

Player segmentation in team sports videos is challenging but crucial to video semantic understanding, such as player interaction identification and tactic analysis. We leverage the appearance similarity among players of the same team, and cast this task as a co-segmentation problem. In this way, the…

Cited by 0SourceScholar
2015

Blur Kernel Estimation Using Normalized Color-Line Prior

CVPR 2015poster

This paper proposes a single-image blur kernel estimation algorithm that utilizes the normalized color-line prior to restore sharp edges without altering edge structures or enhancing noise. The proposed prior is derived from the color-line model, which has been successfully applied to non-blind deco…

Cited by 127SourcePDFScholar
2015

Robust Image Alignment With Multiple Feature Descriptors and Matching-Guided Neighborhoods

CVPR 2015poster

This paper addresses two issues hindering the advances in accurate image alignment. First, the performance of descriptor-based approaches to image alignment relies on the chosen descriptor, but the optimal descriptor typically varies from image to image, or even pixel to pixel. Second, the neighborh…

Cited by 28SourcePDFScholar