← Search

Alex C. Kot

37 accepted papers

2026

ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos

CVPR 2026

Temporal forgery localization aims to temporally identify manipulated segments in videos. Most existing benchmarks focus on appearance-level forgeries, such as face swapping and object removal. However, recent advances in video generation have driven the emergence of activity-level forgeries that mo

Cited by 0SourcecodeScholar
2026

Cross-Domain Few-Shot Segmentation via Multi-view Progressive Adaptation

CVPR 2026

Cross-Domain Few-Shot Segmentation aims to segment categories in data-scarce domains conditioned on a few exemplars. Typical methods first establish few-shot capability in a large-scale source domain and then adapt it to target domains. However, due to the limited quantity and diversity of target sa

Cited by 0SourcecodeScholar
2026

When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models

CVPR 2026

Vision-Language-Action (VLA) models are vulnerable to adversarial attacks, yet universal and transferable attacks remain underexplored, as most existing patches overfit to a single model and fail in black-box settings. To address this gap, we present a systematic study of universal, transferable adv

Cited by 0SourcecodeScholar
2025

Reconciling Stochastic and Deterministic Strategies for Zero-shot Image Restoration using Diffusion Model in Dual

CVPR 2025poster

Plug-and-play (PnP) methods offer an iterative strategy for solving image restoration (IR) problems in a zero-shot manner, using a learned discriminative denoiser as the implicit prior. More recently, a sampling-based variant of this approach, which utilizes a pre-trained generative diffusion model,…

2025

Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object Tracking

ICCV 2025poster

With the rise of social media, vast amounts of user-uploaded videos (e.g., YouTube) are utilized as training data for Visual Object Tracking (VOT). However, the VOT community has largely overlooked video data-privacy issues, as many private videos have been collected and used for training commercial…

Cited by 0SourcePDFScholar
2025

Theoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference Systems

CVPR 2025highlight

By locally encoding raw data into intermediate features, collaborative inference enables end users to leverage powerful deep learning models without exposure of sensitive raw data to cloud servers. However, recent studies have revealed that these intermediate features may not sufficiently preserve p…

2024

Cross-Domain Few-Shot Segmentation via Iterative Support-Query Correspondence Mining

CVPR 2024poster

Cross-Domain Few-Shot Segmentation (CD-FSS) poses the challenge of segmenting novel categories from a distinct domain using only limited exemplars. In this paper we undertake a comprehensive study of CD-FSS and uncover two crucial insights: (i) the necessity of a fine-tuning stage to effectively tra…

2024

Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video Grounding

AAAI 2024technical

This paper for the first time leverages multi-modal videos for weakly-supervised temporal video grounding. As labeling the video moment is labor-intensive and subjective, the weakly-supervised approaches have gained increasing attention in recent years. However, these approaches could inherently com…

Cited by 12SourcePDFScholar
2024

Omnipotent Distillation with LLMs for Weakly-Supervised Natural Language Video Localization: When Divergence Meets Consistency

AAAI 2024technical

Natural language video localization plays a pivotal role in video understanding, and leveraging weakly-labeled data is considered a promising approach to circumvent the laborintensive process of manual annotations. However, this approach encounters two significant challenges: 1) limited input distri…

Cited by 9SourcePDFScholar
2024

SinSR: Diffusion-Based Image Super-Resolution in a Single Step

CVPR 2024poster

While super-resolution (SR) methods based on diffusion models exhibit promising results their practical application is hindered by the substantial number of required inference steps. Recent methods utilize the degraded images in the initial state thereby shortening the Markov chain. Nevertheless the…

2023

Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger

CVPR 2023poster

Recent deep-learning-based compression methods have achieved superior performance compared with traditional approaches. However, deep learning models have proven to be vulnerable to backdoor attacks, where some specific trigger patterns added to the input can lead to malicious behavior of the models…

Cited by 57SourcePDFScholar
2023

Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event Localization

AAAI 2023technical

This paper for the first time explores audio-visual event localization in an unsupervised manner. Previous methods tackle this problem in a supervised setting and require segment-level or video-level event category ground-truth to train the model. However, building large-scale multi-modality dataset…

Cited by 9SourcePDFScholar
2023

ExposureDiffusion: Learning to Expose for Low-light Image Enhancement

ICCV 2023poster

Previous raw image-based low-light image enhancement methods predominantly relied on feed-forward neural networks to learn deterministic mappings from low-light to normally-exposed images. However, they failed to capture critical distribution information, leading to visually undesirable results. Thi…

Cited by 65PDFcodeScholar
2023

Raw Image Reconstruction With Learned Compact Metadata

CVPR 2023poster

While raw images exhibit advantages over sRGB images (e.g. linearity and fine-grained quantization level), they are not widely used by common users due to the large storage requirements. Very recent works propose to compress raw images by designing the sampling masks in the raw image pixel space, le…

2023

Towards Explainable Recommendation Via Bert-Guided Explanation Generator

ICASSP 2023accepted

Explainable recommender system has recently drawn increasing attention due to its capability of providing justification to recommendation. Rather than focusing on certain topics or specific item features, the explanation generated by existing works are too general without the guidance of aspects. Ho…

Cited by 0SourceScholar
2023

Unsupervised Deep Digital Staining for Microscopic Cell Images via Knowledge Distillation

ICASSP 2023accepted

Staining is critical to cell imaging and medical diagnosis, which is expensive, time-consuming, labor-intensive, and causes irreversible changes to cell tissues. Recent advances in deep learning enabled digital staining via supervised model training. However, it is difficult to obtain large-scale st…

Cited by 0SourceScholar
2023

Virtual Try-On with Pose-Garment Keypoints Guided Inpainting

ICCV 2023poster

Virtual try-on is an important technology supporting online apparel shopping, which provides consumers with a virtual experience to fit garments without physically wearing them. Recently, the image-based virtual try-on has received growing research attention. However, the synthetic results of existi…

Cited by 32PDFcodeScholar
2022

Towards Robust Rain Removal Against Adversarial Attacks: A Comprehensive Benchmark Analysis and Beyond

CVPR 2022poster

Rain removal aims to remove rain streaks from images/videos and reduce the disruptive effects caused by rain. It not only enhances image/video visibility but also allows many computer vision algorithms to function properly. This paper makes the first attempt to conduct a comprehensive study on the r…

Cited by 66PDFcodeScholar
2021

Single Image Reflection Removal With Absorption Effect

CVPR 2021poster

In this paper, we consider the absorption effect for the problem of single image reflection removal. We show that the absorption effect can be numerically approximated by the average of refractive amplitude coefficient map. We then reformulate the image formation model and propose a two-step solutio…

Cited by 54PDFcodeScholar
2021

Skeleton Cloud Colorization for Unsupervised 3D Action Representation Learning

ICCV 2021poster

Skeleton-based human action recognition has attracted increasing attention in recent years. However, most of the existing works focus on supervised learning which requiring a large number of annotated action sequences that are often expensive to collect. We investigate unsupervised representation le…

Cited by 123PDFScholar
2020

Collaborative Learning of Gesture Recognition and 3D Hand Pose Estimation with Multi-Order Feature Analysis

ECCV 2020poster

Gesture recognition and 3D hand pose estimation are two highly correlated tasks, yet they are often handled separately. In this paper, we present a novel collaborative learning network for joint gesture recognition and 3D hand pose estimation. The proposed network exploits joint-aware features that…

Cited by 57SourcePDFScholar
2020

What Does Plate Glass Reveal About Camera Calibration?

CVPR 2020poster

This paper aims to calibrate the orientation of glass and the field of view of the camera from a single reflection-contaminated image. We show how a reflective amplitude coefficient map can be used as a calibration cue. Different from existing methods, the proposed solution is free from image conten…

Cited by 19PDFScholar
2019

SPLINE-Net: Sparse Photometric Stereo Through Lighting Interpolation and Normal Estimation Networks

ICCV 2019poster

This paper solves the Sparse Photometric stereo through Lighting Interpolation and Normal Estimation using a generative Network (SPLINE-Net). SPLINE-Net contains a lighting interpolation network to generate dense lighting observations given a sparse set of lights as inputs followed by a normal estim…

Cited by 90PDFScholar
2018

CRRN: Multi-Scale Guided Concurrent Reflection Removal Network

CVPR 2018poster

Removing the undesired reflections from images taken through the glass is of broad application to various computer vision tasks. Non-learning based methods utilize different handcrafted priors such as the separable sparse gradients caused by different levels of blurs, which often fail due to their l…

2018

Domain Generalization With Adversarial Feature Learning

CVPR 2018poster

In this paper, we tackle the problem of domain generalization: how to learn a generalized feature representation for an “unseen” target domain by taking the advantage of multiple seen source-domain data. We present a novel framework based on adversarial autoencoders to learn a generalized latent fea…

Cited by 1574SourcePDFScholar
2018

Dual Attention Matching Network for Context-Aware Feature Sequence Based Person Re-Identification

CVPR 2018poster

Typical person re-identification (ReID) methods usually describe each pedestrian with a single feature vector and match them in a task-specific metric space. However, the methods based on a single feature vector are not sufficient enough to overcome visual ambiguity, which frequently occurs in real…

Cited by 502SourcePDFScholar
2018

SSNet: Scale Selection Network for Online 3D Action Prediction

CVPR 2018poster

In action prediction (early action recognition), the goal is to predict the class label of an ongoing action using its observed part so far. In this paper, we focus on online action prediction in streaming 3D skeleton sequences. A dilated convolutional network is introduced to model the motion dynam…

Cited by 68SourcePDFScholar
2017

Global Context-Aware Attention LSTM Networks for 3D Action Recognition

CVPR 2017poster

Long Short-Term Memory (LSTM) networks have shown superior performance in 3D human action recognition due to their power in modeling the dynamics and dependencies in sequential data. Since not all joints are informative for action analysis and the irrelevant joints often bring a lot of noise, we nee…

Cited by 822PDFScholar