← Search

Robby T. Tan

55 accepted papers

2026

Aggregating Diverse Cue Experts for AI-Generated Image Detection

AAAI 2026technical

The rapid emergence of image synthesis models poses challenges to the generalization of AI-generated image detectors. However, existing methods often rely on model-specific features, leading to overfitting and poor generalization. In this paper, we introduce the Multi-Cue Aggregation Network (MCAN),

Cited by 0SourcePDFScholar
2026

Bridging Day and Night: Target-Class Hallucination Suppression in Unpaired Image Translation

AAAI 2026technical

Day-to-night unpaired image translation is important to downstream tasks but remains challenging due to large appearance shifts and the lack of direct pixel-level supervision. Existing methods often introduce semantic hallucinations, where objects from target classes such as traffic signs and vehicl

Cited by 7SourcePDFScholar
2026

CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods

CVPR 2026

HOI detection has long been dominated by task-specific models, sometimes with early vision-language backbones such as CLIP. With the rise of large generative VLMs, a key question is whether standalone VLMs can perform HOI detection competitively against specialized HOI methods. Existing benchmarks s

Cited by 0SourcecodeScholar
2026

Mind the Gap: Transferring Labels to Align Object Detection Datasets

CVPR 2026

Combining multiple object detection datasets offers a path to improved model generalisation but is hindered by inconsistencies in class semantics and bounding box annotations. Some methods to address this assume shared label taxonomies and address only spatial inconsistencies; others require manual

Cited by 0SourceScholar
2026

UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose Estimation

CVPR 2026

Low-visibility scenarios, such as low-light conditions, pose significant challenges to human pose estimation due to the scarcity of annotated low-light datasets and the loss of visual information under poor illumination. Recent domain adaptation techniques attempt to utilize well-lit labels by augme

Cited by 0SourceScholar
2025

3DOT: Texture Transfer for 3DGS Objects from a Single Reference Image

NeurIPS 2025poster

Image-based 3D texture transfer from a single 2D reference image enables practical customization of 3D object appearances with minimal manual effort. Adapted 2D editing and text-driven 3D editing approaches can serve this purpose. However, 2D editing typically involves frame-by-frame manipulation, o…

Cited by 0SourcecodeScholar
2025

HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation

ICCV 2025poster

Zero-shot human-object interaction (HOI) detection remains a challenging task, particularly in generalizing to unseen actions. Existing methods address this challenge by tapping Vision-Language Models (VLMs) to access knowledge beyond the training data. However, they either struggle to distinguish a…

2025

Mamba as a Bridge: Where Vision Foundation Models Meet Vision Language Models for Domain-Generalized Semantic Segmentation

CVPR 2025highlight

Vision Foundation Models (VFMs) and Vision-Language Models (VLMs) have gained traction in Domain Generalized Semantic Segmentation (DGSS) due to their strong generalization capabilities. However, existing DGSS methods often rely exclusively on either VFMs or VLMs, overlooking their complementary str…

2025

NightHaze: Nighttime Image Dehazing via Self-Prior Learning

AAAI 2025technical

Masked autoencoder (MAE) shows that severe augmentation during training produces robust representations for high-level tasks. This paper brings the MAE-like framework to nighttime image enhancement, demonstrating that severe augmentation during training produces strong network priors that are resili…

Cited by 5SourcePDFScholar
2025

Semantic Segmentation on Raindrop Degraded Images Using Two-Stage Dual Teacher-Student Learning

AAAI 2025technical

Existing semantic segmentation methods face challenges when processing input images degraded by raindrops on the lens or windshield. Unlike other adverse conditions such as fog and nighttime, which degrade visual quality, raindrops not only impair visual appearances but also introduce misleading occ…

Cited by 0SourcePDFScholar
2025

Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure Priors

CVPR 2025poster

Self-supervised depth estimation from monocular cameras in diverse outdoor conditions, such as daytime, rain, and nighttime, is challenging due to the difficulty of learning universal representations and the severe lack of labeled real-world adverse data.Previous methods either rely on synthetic inp…

2025

uMedSum: A Unified Framework for Clinical Abstractive Summarization

ACL 2025long

Clinical abstractive summarization struggles to balance faithfulness and informativeness, sacrificing key information or introducing confabulations. Techniques like in-context learning and fine-tuning have improved overall summary quality orthogonally, without considering the above issue. Conversely…

Cited by 0SourcePDFScholar
2024

Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models

EMNLP 2024finding

Various audio-LLMs (ALLMs) have been explored recently for tackling different audio tasks simultaneously using a single, unified model. While existing evaluations of ALLMs primarily focus on single-audio tasks, real-world applications often involve processing multiple audio streams simultaneously. T…

2024

CAT: Exploiting Inter-Class Dynamics for Domain Adaptive Object Detection

CVPR 2024poster

Domain adaptive object detection aims to adapt detection models to domains where annotated data is unavailable. Existing methods have been proposed to address the domain gap using the semi-supervised student-teacher framework. However a fundamental issue arises from the class imbalance in the labell…

Cited by 11SourcePDFScholar
2024

DeS3: Adaptive Attention-Driven Self and Soft Shadow Removal Using ViT Similarity

AAAI 2024technical

Removing soft and self shadows that lack clear boundaries from a single image is still challenging. Self shadows are shadows that are cast on the object itself. Most existing methods rely on binary shadow masks, without considering the ambiguous boundaries of soft and self shadows. In this paper, we…

2024

Domain-Adaptive 2D Human Pose Estimation via Dual Teachers in Extremely Low-Light Conditions

ECCV 2024poster

"Existing 2D human pose estimation research predominantly concentrates on well-lit scenarios, with limited exploration of poor lighting conditions, which are a prevalent aspect of daily life. Recent studies on low-light pose estimation require the use of paired well-lit and low-light images with gro…

2024

Dual-Rain: Video Rain Removal using Assertive and Gentle Teachers

ECCV 2024poster

"Existing video deraining methods addressing both rain accumulation and rain streaks rely on synthetic data for training as clear ground-truths are unavailable. Hence, they struggle to handle real-world rain videos due to domain gaps. In this paper, we present Dual-Rain, a novel video deraining meth…

Cited by 4SourcePDFScholar
2024

EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI Detection

NeurIPS 2024poster

Detecting Human-Object Interactions (HOI) in zero-shot settings, where models must handle unseen classes, poses significant challenges. Existing methods that rely on aligning visual encoders with large Vision-Language Models (VLMs) to tap into the extensive knowledge of VLMs, require large, computat…

2024

End-to-End Video Semantic Segmentation in Adverse Weather using Fusion Blocks and Temporal-Spatial Teacher-Student Learning

NeurIPS 2024poster

Adverse weather conditions can significantly degrade the video frames, causing existing video semantic segmentation methods to produce erroneous predictions. In this work, we target adverse weather conditions and introduce an end-to-end domain adaptation strategy that leverages a fusion block, tempo…

Cited by 1SourcePDFScholar
2024

HEAP: Unsupervised Object Discovery and Localization with Contrastive Grouping

AAAI 2024technical

Unsupervised object discovery and localization aims to detect or segment objects in an image without any supervision. Recent efforts have demonstrated a notable potential to identify salient foreground objects by utilizing self-supervised transformer features. However, their scopes only build upon p…

Cited by 3SourcePDFScholar
2024

NightRain: Nighttime Video Deraining via Adaptive-Rain-Removal and Adaptive-Correction

AAAI 2024technical

Existing deep-learning-based methods for nighttime video deraining rely on synthetic data due to the absence of real-world paired data. However, the intricacies of the real world, particularly with the presence of light effects and low-light regions affected by noise, create significant domain gaps,…

Cited by 12SourcePDFScholar
2024

Restoring Speaking Lips from Occlusion for Audio-Visual Speech Recognition

AAAI 2024technical

Prior studies on audio-visual speech recognition typically assume the visibility of speaking lips, ignoring the fact that visual occlusion occurs in real-world videos, thus adversely affecting recognition performance. To address this issue, we propose a framework that restores occluded lips in a vid…

Cited by 11SourcePDFScholar
2024

Semantic Segmentation in Multiple Adverse Weather Conditions with Domain Knowledge Retention

AAAI 2024technical

Semantic segmentation's performance is often compromised when applied to unlabeled adverse weather conditions. Unsupervised domain adaptation is a potential approach to enhancing the model's adaptability and robustness to adverse weather. However, existing methods encounter difficulties when sequent…

Cited by 4SourcePDFScholar
2023

2PCNet: Two-Phase Consistency Training for Day-to-Night Unsupervised Domain Adaptive Object Detection

CVPR 2023poster

Object detection at night is a challenging problem due to the absence of night image annotations. Despite several domain adaptation methods, achieving high-precision results remains an issue. False-positive error propagation is still observed in methods using the well-established student-teacher fra…

2023

Auxiliary Tasks Benefit 3D Skeleton-based Human Motion Prediction

ICCV 2023poster

Exploring spatial-temporal dependencies from observed motions is one of the core challenges of human motion prediction. Previous methods mainly focus on dedicated network structures to model the spatial and temporal dependencies. This paper considers a new direction by introducing a model learning f…

Cited by 41PDFcodeScholar
2023

DSFNet: Dual Space Fusion Network for Occlusion-Robust 3D Dense Face Alignment

CVPR 2023poster

Sensitivity to severe occlusion and large view angles limits the usage scenarios of the existing monocular 3D dense face alignment methods. The state-of-the-art 3DMM-based method, directly regresses the model's coefficients, underutilizing the low-level 2D spatial and semantic information, which can…

2023

Deep Homography Mixture for Single Image Rolling Shutter Correction

ICCV 2023poster

We present a deep homography mixture motion model for single image rolling shutter correction. Rolling shutter (RS) effects are often caused by row-wise exposure delay in the widely adopted CMOS sensor. Previous methods often require more than one frame for the correction, leading to data quality re…

Cited by 8PDFcodeScholar
2023

EqMotion: Equivariant Multi-Agent Motion Prediction With Invariant Interaction Reasoning

CVPR 2023poster

Learning to predict agent motions with relationship reasoning is important for many applications. In motion prediction tasks, maintaining motion equivariance under Euclidean geometric transformations and invariance of agent interaction is a critical and fundamental principle. However, such equivaria…

2023

Estimating Reflectance Layer from a Single Image: Integrating Reflectance Guidance and Shadow/Specular Aware Learning

AAAI 2023technical

Estimating the reflectance layer from a single image is a challenging task. It becomes more challenging when the input image contains shadows or specular highlights, which often render an inaccurate estimate of the reflectance layer. Therefore, we propose a two-stage learning method, including refle…

2023

Seeing What You Said: Talking Face Generation Guided by a Lip Reading Expert

CVPR 2023poster

Talking face generation, also known as speech-to-lip generation, reconstructs facial motions concerning lips given coherent speech input. The previous studies revealed the importance of lip-speech synchronization and visual quality. Despite much progress, they hardly focus on the content of lip move…

2022

Unsupervised Night Image Enhancement: When Layer Decomposition Meets Light-Effects Suppression

ECCV 2022poster

"Night images suffer not only from low light, but also from uneven distributions of light. Most existing night visibility enhancement methods focus mainly on enhancing low-light regions. This inevitably leads to over enhancement and saturation in bright regions, such as those regions affected by lig…

2021

DC-ShadowNet: Single-Image Hard and Soft Shadow Removal Using Unsupervised Domain-Classifier Guided Network

ICCV 2021poster

Shadow removal from a single image is generally still an open problem. Most existing learning-based methods use supervised learning and require a large number of paired images (shadow and corresponding non-shadow images) for training. A recent unsupervised method, Mask-ShadowGAN, addresses this limi…

Cited by 152PDFcodeScholar
2021

Graph and Temporal Convolutional Networks for 3D Multi-person Pose Estimation in Monocular Videos

AAAI 2021technical

Despite the recent progress, 3D multi-person pose estimation from monocular videos is still challenging due to the commonly encountered problem of missing information caused by occlusion, partially out-of-frame target persons, and inaccurate person detection. To tackle this problem, we propose a nov…

2021

Monocular 3D Multi-Person Pose Estimation by Integrating Top-Down and Bottom-Up Networks

CVPR 2021poster

In monocular video 3D multi-person pose estimation, inter-person occlusion and close interactions can cause human detection to be erroneous and human-joints grouping to be unreliable. Existing top-down methods rely on human detection and thus suffer from these problems. Existing bottom-up methods do…

Cited by 59PDFcodeScholar
2021

Nighttime Visibility Enhancement by Increasing the Dynamic Range and Suppression of Light Effects

CVPR 2021poster

Most existing nighttime visibility enhancement methods focus on low light. Night images, however, do not only suffer from low light, but also from man-made light effects such as glow, glare, floodlight, etc. Hence, when the existing nighttime visibility enhancement methods are applied to these image…

Cited by 68PDFScholar
2021

Self-Aligned Video Deraining With Transmission-Depth Consistency

CVPR 2021poster

In this paper, we address the problems of rain streaks and rain accumulation removal in video, by developing a self-aligned network with transmission-depth consistency. Existing video based deraining method focus only on rain streak removal, and commonly use optical flow to align the rain video fram…

Cited by 40PDFcodeScholar
2020

Nighttime Defogging Using High-Low Frequency Decomposition and Grayscale-Color Networks

ECCV 2020poster

In this paper, we address the problem of nighttime defogging from a single image. We propose a framework consisting of two main modules: grayscale and color modules. Given an RGB foggy nighttime image, our grayscale module takes the grayscale version of the image as input, and decomposes it into hig…

Cited by 58SourcePDFScholar
2020

Object Tracking using Spatio-Temporal Networks for Future Prediction Location

ECCV 2020poster

We introduce an object tracking algorithm that predicts the future locations of the target object and assists the tracker to handle object occlusion. Given a few frames of an object that are extracted from a complete input sequence, we aim to predict the object’s location in the future frames. To fa…

Cited by 31SourcePDFScholar
2020

Self-Learning Video Rain Streak Removal: When Cyclic Consistency Meets Temporal Correspondence

CVPR 2020poster

In this paper, we address the problem of rain streaks removal in video by developing a self-learned rain streak removal method, which does not require any clean groundtruth images in the training process. The method is inspired by fact that the adjacent frames are highly correlated and can be regard…

Cited by 82PDFcodeScholar
2019

Heavy Rain Image Restoration: Integrating Physics Model and Conditional Adversarial Learning

CVPR 2019poster

Most deraining works focus on rain streaks removal but they cannot deal adequately with heavy rain images. In heavy rain, streaks are strongly visible, dense rain accumulation or rain veiling effect significantly washes out the image, further scenes are relatively more blurry, etc. In this paper, w…

Cited by 451PDFcodeScholar
2019

RainFlow: Optical Flow Under Rain Streaks and Rain Veiling Effect

ICCV 2019poster

Optical flow in heavy rainy scenes is challenging due to the presence of both rain steaks and rain veiling effect, which break the existing optical flow constraints. Concerning this, we propose a deep-learning based optical flow method designed to handle heavy rain. We introduce a feature multiplier…

Cited by 43PDFScholar
2018

Attentive Generative Adversarial Network for Raindrop Removal From a Single Image

CVPR 2018poster

Raindrops adhered to a glass window or camera lens can severely hamper the visibility of a background scene and degrade an image considerably. In this paper, we address the problem by visually removing raindrops, and thus transforming a raindrop degraded image into a clean one. The problem is intrac…

Cited by 851SourcePDFScholar
2017

Deep Joint Rain Detection and Removal From a Single Image

CVPR 2017poster

In this paper, we address a rain removal problem from a single image, even in the presence of heavy rain and rain streak accumulation. Our core ideas lie in our new rain image model and new deep learning architecture. We add a binary map that provides rain streak locations to an existing model, whic…

Cited by 1366PDFScholar
2016

Robust Optical Flow Estimation of Double-Layer Images Under Transparency or Reflection

CVPR 2016poster

This paper deals with a challenging, frequently encountered, yet not properly investigated problem in two-frame optical flow estimation. That is, the input frames are compounds of two imaging layers -- one desired background layer of the scene, and one distracting, possibly moving layer due to trans…

Cited by 62PDFScholar
2015

Simultaneous Video Defogging and Stereo Reconstruction

CVPR 2015poster

We present a method to jointly estimate scene depth and recover the clear latent image from a foggy video sequence. In our formulation, the depth cues from stereo matching and fog information reinforce each other, and produce superior results than conventional stereo or defogging algorithms. We firs…

Cited by 166SourcePDFScholar