← Search

Yap-peng Tan

26 accepted papers

2026

Counterfactual Occlusion-Aware Learning via Visibility Intervention for LiDAR Anomaly Detection

ICML 2026poster

LiDAR point cloud anomaly detection is critical for autonomous system safety, yet most existing methods rely only on visible measurements, overlooking occlusion as a structured consequence of the LiDAR sensing process. We argue that anomalies are characterized not only by what is observed, but also …

Cited by 0SourceScholar
2026

Cross-Domain Few-Shot Segmentation via Multi-view Progressive Adaptation

CVPR 2026

Cross-Domain Few-Shot Segmentation aims to segment categories in data-scarce domains conditioned on a few exemplars. Typical methods first establish few-shot capability in a large-scale source domain and then adapt it to target domains. However, due to the limited quantity and diversity of target sa

Cited by 0SourcecodeScholar
2026

MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation

ICRA 2026poster

Pre-trained Vision-Language-Action (VLA) models have achieved remarkable success in improving robustness and generalization for end-to-end robotic manipulation. However, these models struggle with long-horizon tasks due to their lack of memory and reliance solely on immediate sensory inputs. To addr…

2026

OneHOI: Unifying Human-Object Interaction Generation and Editing

CVPR 2026

Human-Object Interaction (HOI) modelling captures how humans act upon and relate to objects, typically expressed as <person, action, object> triplets. Existing approaches split into two disjoint families: HOI generation synthesises scenes from structured triplets and layout, but fails to integrate m

Cited by 0SourcecodeScholar
2026

Structure-to-Intensity Diffusion for Adverse-Weather LiDAR Generation

CVPR 2026

Adverse-weather LiDAR point cloud generation is challenged by complex weather-induced degradations. These degradations affect geometry and reflectance in fundamentally different ways, making joint modeling difficult and ambiguous, especially when diverse real-world training data is limited. To addre

Cited by 0SourceScholar
2025

Backdoor Attacks Against No-Reference Image Quality Assessment Models via a Scalable Trigger

AAAI 2025technical

No-Reference Image Quality Assessment (NR-IQA), responsible for assessing the quality of a single input image without using any reference, plays a critical role in evaluating and optimizing computer vision systems, e.g., low-light enhancement. Recent research indicates that NR-IQA models are suscep…

2025

CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal Learning

NeurIPS 2025poster

Multimodal machine learning, mimicking the human brain’s ability to integrate various modalities has seen rapid growth. Most previous multimodal models are trained on perfectly paired multimodal input to reach optimal performance. In real‑world deployments, however, the presence of modality is highl…

Cited by 0SourceScholar
2025

MTL-UE: Learning to Learn Nothing for Multi-Task Learning

ICML 2025poster

Most existing unlearnable strategies focus on preventing unauthorized users from training single-task learning (STL) models with personal data. Nevertheless, the paradigm has recently shifted towards multi-task data and multi-task learning (MTL), targeting generalist and foundation models that can h…

2024

Cross-Domain Few-Shot Segmentation via Iterative Support-Query Correspondence Mining

CVPR 2024poster

Cross-Domain Few-Shot Segmentation (CD-FSS) poses the challenge of segmenting novel categories from a distinct domain using only limited exemplars. In this paper we undertake a comprehensive study of CD-FSS and uncover two crucial insights: (i) the necessity of a fine-tuning stage to effectively tra…

2024

InteractDiffusion: Interaction Control in Text-to-Image Diffusion Models

CVPR 2024poster

Large-scale text-to-image (T2I) diffusion models have showcased incredible capabilities in generating coherent images based on textual descriptions enabling vast applications in content generation. While recent advancements have introduced control over factors such as object localization posture and…

2024

Purify Unlearnable Examples via Rate-Constrained Variational Autoencoders

ICML 2024poster

Unlearnable examples (UEs) seek to maximize testing error by making subtle modifications to training examples that are correctly labeled. Defenses against these poisoning attacks can be categorized based on whether specific interventions are adopted during training. The first approach is training-ti…

2023

Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger

CVPR 2023poster

Recent deep-learning-based compression methods have achieved superior performance compared with traditional approaches. However, deep learning models have proven to be vulnerable to backdoor attacks, where some specific trigger patterns added to the input can lead to malicious behavior of the models…

Cited by 57SourcePDFScholar
2023

HOI-aware Adaptive Network for Weakly-supervised Action Segmentation

IJCAI 2023poster

In this paper, we propose an HOI-aware adaptive network named AdaAct for weakly-supervised action segmentation. Most existing methods learn a fixed network to predict the action of each frame with the neighboring frames. However, this would result in ambiguity when estimating similar actions, such a…

Cited by 6SourcePDFScholar
2023

Temporal Coherent Test Time Optimization for Robust Video Classification

ICLR 2023poster

Deep neural networks are likely to fail when the test data is corrupted in real-world deployment (e.g., blur, weather, etc.). Test-time optimization is an effective way that adapts models to generalize to corrupted data during testing, which has been shown in the image domain. However, the technique…

Cited by 17SourcePDFScholar
2022

Learning Transferable Human-Object Interaction Detector With Natural Language Supervision

CVPR 2022poster

It is difficult to construct a data collection including all possible combinations of human actions and interacting objects due to the combinatorial nature of human-object interactions (HOI). In this work, we aim to develop a transferable HOI detector for unseen interactions. Existing HOI detectors…

Cited by 66PDFcodeScholar
2022

Towards Robust Rain Removal Against Adversarial Attacks: A Comprehensive Benchmark Analysis and Beyond

CVPR 2022poster

Rain removal aims to remove rain streaks from images/videos and reduce the disruptive effects caused by rain. It not only enhances image/video visibility but also allows many computer vision algorithms to function properly. This paper makes the first attempt to conduct a comprehensive study on the r…

Cited by 66PDFcodeScholar
2021

Benchmarking the Robustness of Spatial-Temporal Models Against Corruptions

NeurIPS 2021poster

The state-of-the-art deep neural networks are vulnerable to common corruptions (e.g., input data degradations, distortions, and disturbances caused by weather changes, system error, and processing). While much progress has been made in analyzing and improving the robustness of models in image unders…

Cited by 46SourcecodeScholar
2021

Discovering Human Interactions With Large-Vocabulary Objects via Query and Multi-Scale Detection

ICCV 2021poster

In this work, we study the problem of human-object interaction (HOI) detection with large vocabulary object categories. Previous HOI studies are mainly conducted in the regime of limit object categories (e.g., 80 categories). Their solutions may face new difficulties in both object detection and int…

Cited by 32PDFScholar
2020

Discovering Human Interactions With Novel Objects via Zero-Shot Learning

CVPR 2020poster

We aim to detect human interactions with novel objects through zero-shot learning. Different from previous works, we allow unseen object categories by using its semantic word embedding. To do so, we design a human-object region proposal network specifically for the human-object interaction detection…

Cited by 50PDFcodeScholar
2019

Joint Representative Selection and Feature Learning: A Semi-Supervised Approach

CVPR 2019poster

In this paper, we propose a semi-supervised approach for representative selection, which finds a small set of representatives that can well summarize a large data collection. Given labeled source data and big unlabeled target data, we aim to find representatives in the target data, which can not onl…

Cited by 4PDFScholar
2019

Scaling Object Detection by Transferring Classification Weights

ICCV 2019oral

Large scale object detection datasets are constantly increasing their size in terms of the number of classes and annotations count. Yet, the number of object-level categories annotated in detection datasets is an order of magnitude smaller than image-level classification labels. State-of-the art obj…

Cited by 26PDFcodeScholar
2018

Motion-Guided Cascaded Refinement Network for Video Object Segmentation

CVPR 2018poster

Deep CNNs have achieved superior performance in many tasks of computer vision and image understanding. However, it is still difficult to effectively apply deep CNNs to video object segmentation(VOS) since treating video frames as separate and static will lose the information hidden in motion. To tac…

Cited by 131SourcePDFScholar
2018

Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks

CVPR 2018poster

It is desirable to train convolutional networks (CNNs) to run more efficiently during inference. In many cases however, the computational budget that the system has for inference cannot be known beforehand during training, or the inference budget is dependent on the changing real-time resource avail…

Cited by 19SourcePDFScholar
2016

From Keyframes to Key Objects: Video Summarization by Representative Object Proposal Selection

CVPR 2016poster

We propose to summarize a video into a few key objects by selecting representative object proposals generated from video frames. This representative selection problem is formulated as a sparse dictionary selection problem, i.e., choosing a few representatives object proposals to reconstruct the whol…

Cited by 136PDFScholar