← Search

Yu Hen Hu

7 accepted papers

2025

From Prototypes to General Distributions: An Efficient Curriculum for Masked Image Modeling

CVPR 2025poster

Masked Image Modeling (MIM) has emerged as a powerful self-supervised learning paradigm for visual representation learning, enabling models to acquire rich visual representations by predicting masked portions of images from their visible regions. While this approach has shown promising results, we h…

Cited by 0SourcePDFScholar
2023

Why Is Prompt Tuning for Vision-Language Models Robust to Noisy Labels?

ICCV 2023poster

Vision-language models such as CLIP learn a generic text-image embedding from large-scale training data. A vision-language model can be adapted to a new classification task through few-shot prompt tuning. We find that such prompt tuning process is highly robust to label noises. This intrigues us to…

Cited by 19PDFcodeScholar
2021

Efficient Real-Time Video Stabilization with a Novel Least Squares Formulation

ICASSP 2021accepted

We present a novel video stabilization algorithm (LSstab) that removes unwanted motions in real-time. LSstab is based on a novel least squares formulation of the smoothing cost function to alleviate the undesirable camera jitter. A recursive least square solver is derived to minimize the smoothing c…

Cited by 0SourceScholar
2020

Towards Real-Time, Multi-View Video Stereopsis

ICASSP 2020accepted

We present a real-time, multi-view video stereopsis (RTMVS) algorithm. This algorithm processes five synchronized video streams from cameras of a stationary camera array using a commodity laptop computer equipped with an Nvidia GPU. It provides 3D visualization of a dynamic scene from a chosen viewp…

Cited by 0SourceScholar
2018

Adaptive Visual Target Tracking Based on Label Consistent K-Svd Sparse Coding and Kernel Particle Filter

ICASSP 2018accepted

We propose an adaptive visual target tracking algorithm based on Label-Consistent K -Singular Value Decomposition (LC-KSVD) dictionary learning. To construct target templates, local patch features are sampled from foreground and background of the target. LC-KSVD then is applied to these local patche…

Cited by 0SourceScholar