← Search

Hau-San Wong

22 accepted papers

2026

Defect Cue-Preserved Structural Feature Refinement for Few-Shot Anomaly Detection

CVPR 2026

Modern industrial quality control heavily relies on automated anomaly detection. While few-shot anomaly detection addresses the challenge of limited labeled data, real-world inspection faces a vast diversity of anomaly types, sizes, and shapes. We identify the primary cause for the anomaly detection

Cited by 0SourceScholar
2026

Predicting Context-Aware Transcriptional Responses to Unseen Genetic Perturbation Subject to Interactome Distance Constraints

IJCAI 2026

Predicting responses to genetic perturbation is pivotal for elucidating gene regulatory machinery. However, existing methods often rely on statistical perspectives to model differential expression, overlooking the constraints of the underlying molecular interactome, which renders predictions suscept

Cited by 0Scholar
2026

Provable Sample Efficiency of Curriculum Post-Training for Transformer Reasoning

ICML 2026poster

Recent curriculum techniques in the post-training stage of LLMs have been empirically observed to outperform non-curriculum approaches in improving reasoning performance, yet a principled understanding of their effectiveness and limitations remains incomplete. To bridge this gap, we develop an abstr…

Cited by 0SourceScholar
2026

Refinement Contrastive Learning of Cell–Gene Associations for Unsupervised Cell Type Identification

AAAI 2026technical

Unsupervised cell type identification is crucial for uncovering and characterizing heterogeneous populations in single cell omics studies. Although a range of clustering methods have been developed, most focus exclusively on intrinsic cellular structure and ignore the pivotal role of cell-gene assoc

Cited by 0SourcePDFScholar
2026

Syntactic Structure-Guided Visual Grounding with Subject-Centric Feature Enhancement and Verification

IJCAI 2026

Visual grounding aims to localize target objects based on natural language descriptions, and the core challenge lies in the cross-modal gap, which is partly caused by the significant differences in semantic structure between language and vision. Existing methods typically rely on holistic sentence-l

Cited by 0Scholar
2025

Discrete Prior-Based Temporal-Coherent Content Prediction for Blind Face Video Restoration

AAAI 2025technical

Blind face video restoration aims to restore high-fidelity details from videos subjected to complex and unknown degradations. This task poses a significant challenge of managing temporal heterogeneity while at the same time maintaining stable face attributes. In this paper, we introduce a Discrete P…

2025

Prompt-augmented Feature with Cross-domain Contrastive Learning for Efficient Multi-domain Sentiment Analysis

ICASSP 2025accepted

Pre-trained language models (PrLMs) demonstrate impressive performance on the sentiment analysis task. However, the large number of trainable parameters brings about heavy computational costs, which become more serious in multi-domain scenarios. In this paper, we propose to extract multi-layer featu…

Cited by 0SourceScholar
2025

Provable In-Context Vector Arithmetic via Retrieving Task Concepts

ICML 2025poster

In-context learning (ICL) has garnered significant attention for its ability to grasp functions/tasks from demonstrations. Recent studies suggest the presence of a latent **task/function vector** in LLMs during ICL. Merullo et al. (2024) showed that LLMs leverage this vector alongside the residual s…

Cited by 0SourcePDFScholar
2025

RetouchGPT: LLM-based Interactive High-Fidelity Face Retouching via Imperfection Prompting

AAAI 2025technical

Face retouching aims to remove facial imperfections from image and videos while at the same time preserving face attributes. The existing methods are designed to perform non-interactive end-to-end retouching, while the ability to interact with users is highly demanded in downstream applications. In…

Cited by 0SourcePDFScholar
2025

SpotDiff: Spatial Gene Expression Imputation Diffusion with Single-Cell RNA Sequencing Data Integration

AAAI 2025technical

The advent of Spatial Transcriptomics (ST) has revolutionized understanding of tissue architecture by creating high-resolution maps of gene expression patterns. However, the low capture rate of ST leads to significant sparsity. The aim of imputation is to recover biological signals by imputing the d…

Cited by 0SourcePDFScholar
2025

Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual Grounding

CVPR 2025poster

The goal of visual grounding is to establish connections between target objects and textual descriptions. Large Language Models (LLMs) have demonstrated strong comprehension abilities across a variety of visual tasks. To establish precise associations between the text and the corresponding visual re…

Cited by 0SourcePDFScholar
2024

Provably Neural Active Learning Succeeds via Prioritizing Perplexing Samples

ICML 2024poster

Neural Network-based active learning (NAL) is a cost-effective data selection technique that utilizes neural networks to select and train on a small subset of samples. While existing work successfully develops various effective or theory-justified NAL algorithms, the understanding of the two commonl…

Cited by 3SourcePDFScholar
2024

Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning

NeurIPS 2024poster

Transformer-based large language models (LLMs) have displayed remarkable creative prowess and emergence capabilities. Existing empirical studies have revealed a strong connection between these LLMs' impressive emergence abilities and their in-context learning (ICL) capacity, allowing them to solve n…

Cited by 0SourcePDFScholar
2024

Relational Matching for Weakly Semi-Supervised Oriented Object Detection

CVPR 2024poster

Oriented object detection has witnessed significant progress in recent years. However the impressive performance of oriented object detectors is at the huge cost of labor-intensive annotations and deteriorates once the annotated data becomes limited. Semi-supervised learning in which sufficient unan…

Cited by 4SourcePDFScholar
2024

RetouchFormer: Semi-supervised High-Quality Face Retouching Transformer with Prior-Based Selective Self-Attention

AAAI 2024technical

Face retouching is to beautify a face image, while preserving the image content as much as possible. It is a promising yet challenging task to remove face imperfections and fill with normal skin. Generic image enhancement methods are hampered by the lack of imperfection localization, which often res…

Cited by 1SourcePDFScholar
2021

High Fidelity GAN Inversion via Prior Multi-Subspace Feature Composition

AAAI 2021technical

Generative Adversarial Networks (GANs) have shown impressive gains in image synthesis. GAN inversion was recently studied to understand and utilize the knowledge it learns, where a real image is inverted back to a latent code and can thus be reconstructed by the generator. Although increasing the nu…

Cited by 0SourcePDFScholar
2021

Mask-Embedded Discriminator With Region-Based Semantic Regularization for Semi-Supervised Class-Conditional Image Synthesis

CVPR 2021poster

Semi-supervised generative learning (SSGL) makes use of unlabeled data to achieve a trade-off between the data collection/annotation effort and generation performance, when adequate labeled data are not available. Learning precise class semantics is crucial for class-conditional image synthesis with…

Cited by 6PDFScholar
2020

Model Adaptation: Unsupervised Domain Adaptation Without Source Data

CVPR 2020poster

In this paper, we investigate a challenging unsupervised domain adaptation setting --- unsupervised model adaptation. We aim to explore how to rely only on unlabeled target data to improve performance of an existing source prediction model on the target domain, since labeled source data may not be a…

Cited by 0PDFScholar
2020

Regularizing Discriminative Capability of CGANs for Semi-Supervised Generative Learning

CVPR 2020poster

Semi-supervised generative learning aims to learn the underlying class-conditional distribution of partially labeled data. Generative Adversarial Networks (GANs) have led to promising progress in this task. However, it still needs to further explore the issue of imbalance between real labeled data a…

Cited by 29PDFScholar
2019

Enhancing TripleGAN for Semi-Supervised Conditional Instance Synthesis and Classification

CVPR 2019poster

Learning class-conditional data distributions is crucial for Generative Adversarial Networks (GAN) in semi-supervised learning. To improve both instance synthesis and classification in this setting, we propose an enhanced TripleGAN (EnhancedTGAN) model in this work. We follow the adversarial trainin…

Cited by 41PDFScholar
2019

Mutual Learning of Complementary Networks via Residual Correction for Improving Semi-Supervised Classification

CVPR 2019oral

Deep mutual learning jointly trains multiple essential networks having similar properties to improve semi-supervised classification. However, the commonly used consistency regularization between the outputs of the networks may not fully leverage the difference between them. In this paper, we explore…

Cited by 44PDFScholar
2019

Semi-Supervised Pedestrian Instance Synthesis and Detection With Mutual Reinforcement

ICCV 2019poster

We propose a GAN-based scene-specific instance synthesis and classification model for semi-supervised pedestrian detection. Instead of collecting unreliable detections from unlabeled data, we adopt a class-conditional GAN for synthesizing pedestrian instances to alleviate the problem of insufficient…

Cited by 12PDFScholar