← Search

Chien-Yi Wang

16 accepted papers

2026

Composite-Attribute Person Re-Identification via Pose-Guided Disentanglement

CVPR 2026

Recent advancements in vision-language models have enabled multi-modal person re-identification (Re-ID), where the system takes both an image and a text query to identify matching individuals. While previous state-of-the-art methods perform well with detailed, sentence-level descriptions, we found t

Cited by 0SourceScholar
2026

Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos

CVPR 2026

Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing approaches assume specific underlying physical systems, object types, and camera poses, which are unable to generalize to complex real-world settin

Cited by 0SourceScholar
2026

V2V-GoT: Vehicle-To-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-Of-Thoughts

ICRA 2026poster

Current state-of-the-art autonomous vehicles could face safety critical situations when their local sensors are occluded by large objects on the road nearby. Vehicle-to-vehicle (V2V) cooperative autonomous driving is proposed to address this problem. More recent work further adopts a new approach th…

2026

V2V-LLM: Vehicle-To-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models

ICRA 2026poster

Current autonomous driving vehicles rely mainly on their individual sensors to understand surrounding scenes and plan for future trajectories, which can be unreliable when the sensors are malfunctioning or occluded. To address this problem, cooperative perception methods via vehicle-to-vehicle (V2V)…

2025

BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL

ICLR 2025poster

Bayesian optimization (BO) offers an efficient pipeline for optimizing black-box functions with the help of a Gaussian process prior and an acquisition function (AF). Recently, in the context of single-objective BO, learning-based AFs witnessed promising empirical results given its favorable non-myo…

Cited by 0SourcePDFScholar
2025

SANER: Annotation-free Societal Attribute Neutralizer for Debiasing CLIP

ICLR 2025poster

Large-scale vision-language models, such as CLIP, are known to contain societal bias regarding protected attributes (e.g., gender, age). This paper aims to address the problems of societal bias in CLIP. Although previous studies have proposed to debias societal bias through adversarial learning or t…

Cited by 2SourcePDFScholar
2024

DoRA: Weight-Decomposed Low-Rank Adaptation

ICML 2024oral

Among the widely used parameter-efficient fine-tuning (PEFT) methods, LoRA and its variants have gained considerable popularity because of avoiding additional inference costs. However, there still often exists an accuracy gap between these methods and full fine-tuning (FT). In this work, we first in…

2024

MCPNet: An Interpretable Classifier via Multi-Level Concept Prototypes

CVPR 2024poster

Recent advancements in post-hoc and inherently interpretable methods have markedly enhanced the explanations of black box classifier models. These methods operate either through post-analysis or by integrating concept learning during model training. Although being effective in bridging the semantic…

2024

Probabilistic 3D Multi-Object Cooperative Tracking for Autonomous Driving via Differentiable Multi-Sensor Kalman Filter

ICRA 2024poster

Current state-of-the-art autonomous driving vehicles mainly rely on each individual sensor system to perform perception tasks. Such a framework’s reliability could be limited by occlusion or sensor failure. To address this issue, more recent research proposes using vehicle-to-vehicle (V2V) communica…

Cited by 8SourcecodeScholar
2024

RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question Answering

ICLR 2024poster

Natural Language Explanation (NLE) in vision and language tasks aims to provide human-understandable explanations for the associated decision-making process. In practice, one might encounter explanations which lack informativeness or contradict visual-grounded facts, known as implausibility and hall…

Cited by 2SourcePDFScholar
2023

Efficient Model Personalization in Federated Learning via Client-Specific Prompt Generation

ICCV 2023poster

Federated learning (FL) emerges as a decentralized learning framework which trains models from multiple distributed clients without sharing their data to preserve privacy. Recently, large-scale pre-trained models (e.g., Vision Transformer) have shown a strong capability of deriving robust representa…

Cited by 46PDFScholar
2023

MixFairFace: Towards Ultimate Fairness via MixFair Adapter in Face Recognition

AAAI 2023technical

Although significant progress has been made in face recognition, demographic bias still exists in face recognition systems. For instance, it usually happens that the face recognition performance for a certain demographic group is lower than the others. In this paper, we propose MixFairFace framework…

2022

FedFR: Joint Optimization Federated Framework for Generic and Personalized Face Recognition

AAAI 2022technical

Current state-of-the-art deep learning based face recognition (FR) models require a large number of face identities for central training. However, due to the growing privacy awareness, it is prohibited to access the face images on user devices to continually improve face recognition models. Federate…

2022

Local-Adaptive Face Recognition via Graph-Based Meta-Clustering and Regularized Adaptation

CVPR 2022poster

Due to the rising concern of data privacy, it's reasonable to assume the local client data can't be transferred to a centralized server, nor their associated identity label is provided. To support continuous learning and fill the last-mile quality gap, we introduce a new problem setup called Local-A…

Cited by 14PDFScholar
2022

PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition

CVPR 2022poster

Face anti-spoofing (FAS) plays a critical role in securing face recognition systems from different presentation attacks. Previous works leverage auxiliary pixel-level supervision and domain generalization approaches to address unseen spoof types. However, the local characteristics of image captures,…

Cited by 147PDFScholar
2015

Robust Image Segmentation Using Contour-Guided Color Palettes

ICCV 2015poster

The contour-guided color palette (CCP) is proposed for robust image segmentation. It efficiently integrates contour and color cues of an image. To find representative colors of an image, color samples along long contours between regions, similar in spirit to machine learning methodology that focus o…

Cited by 35PDFcodeScholar