← Search

Wen-Sheng Chu

17 accepted papers

2026

Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification

CVPR 2026

Conventional evaluation methods for multimodal LLMs (MLLMs) lack interpretability and are often insufficient to fully disclose significant capability gaps across models. To address this, we introduce AuditDM, an automated framework that actively discovers and rectifies MLLM failure modes by auditing

Cited by 0SourceScholar
2026

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception

ICLR 2026poster

Vision foundation models are typically trained as static feature extractors, forcing the burden of task adaptation onto large downstream models. We propose a different paradigm: instead of solely feeding visual features into language, we use language itself to dynamically guide the vision encoder. O…

Cited by 0SourceScholar
2026

Mining Attribute Subspaces for Efficient Fine-tuning of 3D Foundation Models

CVPR 2026

With the emergence of 3D foundation models, there is growing interest in fine-tuning them for downstream tasks, where LoRA is the dominant fine-tuning paradigm. As 3D datasets exhibit distinct variations in texture, geometry, camera motion, and lighting, there are interesting fundamental questions:

Cited by 0SourceScholar
2025

RB-Modulation: Training-Free Stylization using Reference-Based Modulation

ICLR 2025oral

We propose Reference-Based Modulation (RB-Modulation), a new plug-and-play solution for training-free personalization of diffusion models. Existing training-free approaches exhibit difficulties in (a) style extraction from reference images in the absence of additional style or content text descripti…

2025

SEAL: Semantic Attention Learning for Long Video Representation

CVPR 2025poster

Long video understanding presents challenges due to the inherent high computational complexity and redundant temporal information. An effective representation for long videos must efficiently process such redundancy while preserving essential contents for downstream tasks. This paper introduces **S…

Cited by 0SourcePDFScholar
2025

Semantic Image Inversion and Editing using Rectified Stochastic Differential Equations

ICLR 2025poster

Generative models transform random noise into images, while their inversion aims to reconstruct structured noise for recovery and editing. This paper addresses two key tasks: (i) *inversion* and (ii) *editing* of real images using stochastic equivalents of rectified flow models (e.g., Flux). While D…

2024

Beyond First-Order Tweedie: Solving Inverse Problems using Latent Diffusion

CVPR 2024poster

Sampling from the posterior distribution in latent diffusion models for inverse problems is computationally challenging. Existing methods often rely on Tweedie's first-order moments that tend to induce biased results. Second-order approximations are computationally prohibitive making standard revers…

Cited by 26SourcePDFScholar
2023

Distributionally Robust Post-hoc Classifiers under Prior Shifts

ICLR 2023poster

The generalization ability of machine learning models degrades significantly when the test distribution shifts away from the training distribution. We investigate the problem of training models that are robust to shifts caused by changes in the distribution of class-priors or group-priors. The prese…

2023

Rethinking Domain Generalization for Face Anti-Spoofing: Separability and Alignment

CVPR 2023poster

This work studies the generalization issue of face anti-spoofing (FAS) models on domain gaps, such as image resolution, blurriness and sensor variations. Most prior works regard domain-specific signals as a negative impact, and apply metric learning or adversarial losses to remove it from feature re…

2022

Adaptive Transformers for Robust Few-Shot Cross-Domain Face Anti-Spoofing

ECCV 2022poster

"While recent face anti-spoofing methods perform well under the intra-domain setups, an effective approach needs to account for much larger appearance variations of images acquired in complex scenes with different sensors for robust performance. In this paper, we present adaptive vision transformers…

Cited by 95SourcePDFScholar
2021

Retrieve in Style: Unsupervised Facial Feature Transfer and Retrieval

ICCV 2021poster

We present Retrieve in Style (RIS), an unsupervised framework for facial feature transfer and retrieval on real images. Recent work shows capabilities of transferring local facial features by capitalizing on the disentanglement property of the StyleGAN latent space. RIS improves existing art on the…

Cited by 31PDFcodeScholar
2018

Learning Facial Action Units From Web Images With Scalable Weakly Supervised Clustering

CVPR 2018poster

We present a scalable weakly supervised clustering approach to learn facial action units (AUs) from large, freely available web images. Unlike most existing methods (e.g., CNNs) that rely on fully annotated data, our method exploits web images with inaccurate annotations. Specifically, we derive a w…

Cited by 71SourcePDFScholar
2015

Confidence Preserving Machine for Facial Action Unit Detection

ICCV 2015poster

Varied sources of error contribute to the challenge of facial action unit detection. Previous approaches address specific and known sources. However, many sources are unknown. To address the ubiquity of error, we propose a Confident Preserving Machine (CPM) that follows an easy-to-hard classificatio…

Cited by 82PDFScholar
2015

Joint Patch and Multi-Label Learning for Facial Action Unit Detection

CVPR 2015poster

The face is one of the most powerful channel of non-verbal communication. The most commonly used taxonomy to describe facial behaviour is the Facial Action Coding System (FACS). FACS segments the visible effects of facial muscle activation into 30+ action units (AUs). AUs, which may occur alone an…

Cited by 246SourcePDFScholar
2015

Unsupervised Synchrony Discovery in Human Interaction

ICCV 2015poster

People are inherently social. Social interaction plays an important and natural role in human behavior. Most computational methods focus on individuals alone rather than in social context. They also require labelled training data. We present an unsupervised approach to discover interpersonal synchro…

Cited by 19PDFScholar