← Search

Alex Kot

22 accepted papers

2026

From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge

AAAI 2026technical

Large-scale Video Foundation Models (VFMs) have significantly advanced various video-related tasks, either through task-specific models or Multi-modal Large Language Models (MLLMs). However, the open accessibility of VFMs also introduces critical security risks, as adversaries can exploit full knowl

Cited by 0SourcePDFScholar
2026

Random Erasing vs. Model Inversion: A Promising Defense or a False Hope?

ICML 2026poster

Model Inversion (MI) attacks pose a significant privacy threat by reconstructing private training data from machine learning models. While existing defenses primarily concentrate on model-centric approaches, the impact of data on MI robustness remains largely unexplored. In this work, we explore Ran…

Cited by 0SourcecodeScholar
2026

SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision

AAAI 2026technical

Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previous mitigation methods effectively reduce hallucinations in photographic images, t

Cited by 0SourcePDFScholar
2025

Backdoor Attacks Against No-Reference Image Quality Assessment Models via a Scalable Trigger

AAAI 2025technical

No-Reference Image Quality Assessment (NR-IQA), responsible for assessing the quality of a single input image without using any reference, plays a critical role in evaluating and optimizing computer vision systems, e.g., low-light enhancement. Recent research indicates that NR-IQA models are suscep…

2025

MTL-UE: Learning to Learn Nothing for Multi-Task Learning

ICML 2025poster

Most existing unlearnable strategies focus on preventing unauthorized users from training single-task learning (STL) models with personal data. Nevertheless, the paradigm has recently shifted towards multi-task data and multi-task learning (MTL), targeting generalist and foundation models that can h…

2025

Vid-Group: Temporal Video Grounding Pretraining from Unlabeled Videos in the Wild

ICCV 2025poster

Given a natural language query, temporal video grounding aims to localize the described temporal moment in an untrimmed video. A major challenge of this task is its heavy dependence on labor-intensive annotations for training. Unlike existing works that directly train models on manually curated data…

2024

BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models

ECCV 2024poster

"Large Multimodal Models (LMMs) such as GPT-4V and LLaVA have shown remarkable capabilities in visual reasoning on data in common image styles. However, their robustness against diverse style shifts, crucial for practical applications, remains largely unexplored. In this paper, we propose a new benc…

2024

Compress Clean Signal from Noisy Raw Image: A Self-Supervised Approach

ICML 2024poster

Raw images offer unique advantages in many low-level visual tasks due to their unprocessed nature. However, this unprocessed state accentuates noise, making raw images challenging to compress effectively. Current compression methods often overlook the ubiquitous noise in raw space, leading to increa…

Cited by 0SourcePDFScholar
2024

ContextGS : Compact 3D Gaussian Splatting with Anchor Level Context Model

NeurIPS 2024poster

Recently, 3D Gaussian Splatting (3DGS) has become a promising framework for novel view synthesis, offering fast rendering speeds and high fidelity. However, the large number of Gaussians and their associated attributes require effective compression techniques. Existing methods primarily compress ne…

2024

E3M: Zero-Shot Spatio-Temporal Video Grounding with Expectation-Maximization Multimodal Modulation

ECCV 2024oral

"Spatio-temporal video grounding aims to localize the spatio-temporal tube in a video according to the given language query. To eliminate the annotation costs, we make a first exploration to tackle spatio-temporal video grounding in a zero-shot manner. Our method dispenses with the need for any trai…

2024

Purify Unlearnable Examples via Rate-Constrained Variational Autoencoders

ICML 2024poster

Unlearnable examples (UEs) seek to maximize testing error by making subtle modifications to training examples that are correctly labeled. Defenses against these poisoning attacks can be categorized based on whether specific interventions are adopted during training. The first approach is training-ti…

2024

STSP: Spatial-Temporal Subspace Projection for Video Class-incremental Learning

ECCV 2024poster

"Video class-incremental learning (VCIL) aims to learn discriminative and generalized feature representations for video frames to mitigate catastrophic forgetting. Conventional VCIL methods often retain a subset of frames or features from prior tasks as exemplars for subsequent incremental learning…

Cited by 3SourcePDFScholar
2024

Towards Physical World Backdoor Attacks against Skeleton Action Recognition

ECCV 2024poster

"Skeleton Action Recognition (SAR) has attracted significant interest for its efficient representation of the human skeletal structure. Despite its advancements, recent studies have raised security concerns in SAR models, particularly their vulnerability to adversarial attacks. However, such strateg…

Cited by 3SourcePDFScholar
2023

Audio-Visual Deception Detection: DOLOS Dataset and Parameter-Efficient Crossmodal Learning

ICCV 2023poster

Deception detection in conversations is a challenging yet important task, having pivotal applications in many fields such as credibility assessment in business, multimedia anti-frauds, and custom security. Despite this, deception detection research is hindered by the lack of high-quality deception d…

Cited by 12PDFcodeScholar
2023

Rehearsal-Free Domain Continual Face Anti-Spoofing: Generalize More and Forget Less

ICCV 2023oral

Face Anti-Spoofing (FAS) is recently studied under the continual learning setting, where the FAS models are expected to evolve after encountering data from new domains. However, existing methods need extra replay buffers to store previous data for rehearsal, which becomes infeasible when previous da…

Cited by 25PDFcodeScholar
2023

Temporal Coherent Test Time Optimization for Robust Video Classification

ICLR 2023poster

Deep neural networks are likely to fail when the test data is corrupted in real-world deployment (e.g., blur, weather, etc.). Test-time optimization is an effective way that adapts models to generalize to corrupted data during testing, which has been shown in the image domain. However, the technique…

Cited by 17SourcePDFScholar
2022

Low-Light Image Enhancement with Normalizing Flow

AAAI 2022technical

To enhance low-light images to normally-exposed ones is highly ill-posed, namely that the mapping relationship between them is one-to-many. Previous works based on the pixel-wise reconstruction losses and deterministic processes fail to capture the complex conditional distribution of normally expose…

2021

Benchmarking the Robustness of Spatial-Temporal Models Against Corruptions

NeurIPS 2021poster

The state-of-the-art deep neural networks are vulnerable to common corruptions (e.g., input data degradations, distortions, and disturbances caused by weather changes, system error, and processing). While much progress has been made in analyzing and improving the robustness of models in image unders…

Cited by 46SourcecodeScholar
2020

Domain Generalization for Medical Imaging Classification with Linear-Dependency Regularization

NeurIPS 2020poster

Recently, we have witnessed great progress in the field of medical imaging classification by adopting deep neural networks. However, the recent advanced models still require accessing sufficiently large and representative datasets for training, which is often unfeasible in clinically realistic envir…

2020

Splitting vs. Merging: Mining Object Regions with Discrepancy and Intersection Loss for Weakly Supervised Semantic Segmentation

ECCV 2020poster

In this paper we focus on the task of weakly-supervised semantic segmentation supervised with image-level labels. Since the pixel-level annotation is not available in the training process, we rely on region mining models to estimate the pseudo-masks from the image-level labels. Thus, in order to imp…

Cited by 81SourcePDFScholar