← Search

Shang-Hong Lai

18 accepted papers

2026

TIPO: Text to Image with Text Pre-sampling for Prompt Optimization

ICLR 2026poster

TIPO (Text-to-Image Prompt Optimization) introduces an efficient approach for automatic prompt refinement in text-to-image (T2I) generation. Starting from simple user prompts, TIPO leverages a lightweight pre-trained model to expand these prompts into richer, detailed versions. Conceptually, TIPO sa…

Cited by 0SourceScholar
2025

HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics

ICCV 2025poster

Long-form video understanding presents unique challenges that extend beyond traditional short-video analysis approaches, particularly in capturing long-range dependencies, processing redundant information efficiently, and extracting high-level semantic concepts. To address these challenges, we propo…

2025

InstAD: Instance-aware Segmentation Framework for Zero-shot Multi-instance Anomaly Detection

ICASSP 2025accepted

Multi-instance anomaly detection and segmentation play a crucial role in automated industrial inspection. Previous works mainly focus on single-instance detection tasks that require well-aligned input and extensive training sets. In this work, we introduce InstAD, a zero-shot multi-instance anomaly…

Cited by 0SourceScholar
2025

KeyGS: A Keyframe-Centric Gaussian Splatting Method for Monocular Image Sequences

AAAI 2025technical

Reconstructing high-quality 3D models from sparse 2D images has garnered significant attention in computer vision. Recently, 3D Gaussian Splatting (3DGS) has gained prominence due to its explicit representation with efficient training speed and real-time rendering capabilities. However, existing met…

Cited by 0SourcePDFScholar
2025

MovieCORE: COgnitive REasoning in Movies

EMNLP 2025

This paper introduces MovieCORE, a novel video question answering (VQA) dataset designed to probe deeper cognitive understanding of movie content. Unlike existing datasets that focus on surface-level comprehension, MovieCORE emphasizes questions that engage System-2 thinking while remaining specific

2023

MixFairFace: Towards Ultimate Fairness via MixFair Adapter in Face Recognition

AAAI 2023technical

Although significant progress has been made in face recognition, demographic bias still exists in face recognition systems. For instance, it usually happens that the face recognition performance for a certain demographic group is lower than the others. In this paper, we propose MixFairFace framework…

2023

ReST: A Reconfigurable Spatial-Temporal Graph Model for Multi-Camera Multi-Object Tracking

ICCV 2023poster

Multi-Camera Multi-Object Tracking (MC-MOT) utilizes information from multiple views to better handle problems with occlusion and crowded scenes. Recently, the use of graph-based approaches to solve tracking problems has become very popular. However, many current graph-based methods do not effective…

Cited by 28PDFcodeScholar
2022

FedFR: Joint Optimization Federated Framework for Generic and Personalized Face Recognition

AAAI 2022technical

Current state-of-the-art deep learning based face recognition (FR) models require a large number of face identities for central training. However, due to the growing privacy awareness, it is prohibited to access the face images on user devices to continually improve face recognition models. Federate…

2022

Local-Adaptive Face Recognition via Graph-Based Meta-Clustering and Regularized Adaptation

CVPR 2022poster

Due to the rising concern of data privacy, it's reasonable to assume the local client data can't be transferred to a centralized server, nor their associated identity label is provided. To support continuous learning and fill the last-mile quality gap, we introduce a new problem setup called Local-A…

Cited by 14PDFScholar
2022

PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition

CVPR 2022poster

Face anti-spoofing (FAS) plays a critical role in securing face recognition systems from different presentation attacks. Previous works leverage auxiliary pixel-level supervision and domain generalization approaches to address unseen spoof types. However, the local characteristics of image captures,…

Cited by 147PDFScholar
2018

AugGAN: Cross Domain Adaptation with GAN-based Data Augmentation

ECCV 2018poster

Deep learning based image-to-image translation methods aim at learning the joint distribution of the two domains and finding transformations between them. Despite recent GAN (Generative Adversarial Network) based methods have shown compelling visual results, they are prone to fail at preserving imag…

Cited by 350SourcePDFScholar
2017

A deep learning approach towards pore extraction for high-resolution fingerprint recognition

ICASSP 2017accepted

As high-resolution fingerprint images are becoming more common, the pores have been found to be one of the promising candidates in improving the performance of automated fingerprint identification systems (AFIS). This paper proposes a deep learning approach towards pore extraction. It exploits the f…

Cited by 0SourceScholar
2017

Hierarchical Structured Dictionary Learning for image categorization

ICASSP 2017accepted

A novel Hierarchical Structured Dictionary Learning (HSDL) algorithm is proposed in this paper. It aims to learn class-specific dictionaries for all classes simultaneously in a hierarchical structure. A discriminative term based on Fisher discrimination criterion is jointly considered for both the c…

Cited by 0SourceScholar
2015

Non-Rigid Registration of Images With Geometric and Photometric Deformation by Using Local Affine Fourier-Moment Matching

CVPR 2015poster

Registration between images taken with different cameras, from different viewpoints or under different lighting conditions is a challenging problem. It needs to solve not only the geometric registration problem but also the photometric matching problem. In this paper, we propose to estimate the inte…

Cited by 31SourcePDFScholar