← Search

Shu-Tao Xia

148 accepted papers

2026

CASL: Curvature-Augmented Self-supervised Learning for 3D Anomaly Detection

AAAI 2026technical

Deep learning-based 3D anomaly detection methods have demonstrated significant potential in industrial manufacturing. However, many approaches are specifically designed for anomaly detection tasks, which limits their generalizability to other 3D tasks. In contrast, self-supervised point cloud models

Cited by 0SourcePDFScholar
2026

Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive Models

ICLR 2026poster

The rapid progress of visual autoregressive (VAR) models has brought new opportunities for text-to-image generation, but also heightened safety concerns. Existing concept erasure techniques, primarily designed for diffusion models, fail to generalize to VARs due to their next-scale token prediction…

Cited by 0SourcecodeScholar
2026

FreqSIC: Frequency-aware Stereo Image Compression with Bi-directional Checkerboard Context Model

CVPR 2026

Stereo image compression is essential for a wide range of 3D vision. Recent methods have demonstrated strong capabilities in eliminating inter-view redundancy and enabling compact entropy coding via spatial-domain stereo transformation and advanced autoregressive entropy models. However, these appro

Cited by 0SourceScholar
2026

Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval

CVPR 2026

Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos based on text queries that describe only partial events. Existing methods suffer from incomplete global contextual perception, struggling with query ambiguity and local noise induced by spurious responses. To address these i

Cited by 0SourcecodeScholar
2026

Improving Deepfake Detection with Reinforcement Learning-Based Adaptive Data Augmentation

AAAI 2026technical

The generalization capability of deepfake detectors is crucial for real-world applications. Data augmentation to generate synthetic fake faces has served as an effective strategy to enhance generalization. Interestingly, current state-of-the-art (SoTA) methods rely on fixed augmentation strategies,

Cited by 0SourcePDFScholar
2026

Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning

CVPR 2026

Visual in-context learning (VICL) enables visual foundation models to handle multiple tasks by steering them with demonstrative prompts. The choice of such prompts largely influences VICL performance, standing out as a key challenge. Prior work has made substantial progress on prompt retrieval and r

Cited by 0SourcecodeScholar
2026

MambaSIC: Mamba-based Stereo Image Compression with Bi-directional Multi-reference Entropy Model

CVPR 2026

Stereo image compression (SIC) has become increasingly vital with its applications surging in fields such as 3D reconstruction and autonomous navigation. Previous methods leverage cross-attention to model inter-view redundancy and employ autoregressive entropy models to predict probability distribut

Cited by 0SourceScholar
2026

NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale

ICLR 2026oral

Prevailing autoregressive (AR) models for text-to-image generation either rely on heavy, computationally-intensive diffusion models to process continuous image tokens, or employ vector quantization (VQ) to obtain discrete tokens with quantization loss. In this paper, we push the autoregressive parad…

Cited by 0SourcecodeScholar
2026

PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and Alignment

ICLR 2026poster

Visual In-Context Learning (VICL) aims to complete vision tasks by imitating pixel demonstrations. Recent work Condenser pioneered prompt fusion that combines the advantages of various demonstrations, which shows a promising way to extend VICL. Unfortunately, the patch-wise fusion framework and mode…

Cited by 0SourceScholar
2026

RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization

ICLR 2026poster

Visual manipulation localization (VML) aims to identify tampered regions in images and videos, a task that has become increasingly challenging with the rise of advanced editing tools. Existing methods face two main issues: resolution diversity, where resizing or padding distorts forensic traces and…

Cited by 0SourcecodeScholar
2026

SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration

ICLR 2026poster

Vision-Language-Action (VLA) models have attracted increasing attention for their strong control capabilities. However, their high computational cost and low execution frequency hinder their suitability for real-time tasks such as robotic manipulation and autonomous navigation. Existing VLA accelera…

Cited by 0SourcecodeScholar
2026

SparseEval: Efficient Evaluation of Large Language Models by Sparse Optimization

ICLR 2026poster

As large language models (LLMs) continue to scale up, their performance on various downstream tasks has significantly improved. However, evaluating their capabilities has become increasingly expensive, as performing inference on a large number of benchmark samples incurs high computational costs. In…

Cited by 0SourcecodeScholar
2026

VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping

CVPR 2026

Visual autoregressive (AR) generation models have demonstrated strong potential for image generation, yet their next-token-prediction paradigm introduces considerable inference latency. Although speculative decoding (SD) has been proven effective for accelerating visual AR models, its "draft one ste

Cited by 0SourcecodeScholar
2025

3D-LMVIC: Learning-based Multi-View Image Compression with 3D Gaussian Geometric Priors

ICML 2025poster

Existing multi-view image compression methods often rely on 2D projection-based similarities between views to estimate disparities. While effective for small disparities, such as those in stereo images, these methods struggle with the more complex disparities encountered in wide-baseline multi-camer…

Cited by 0SourcePDFScholar
2025

Adapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal Perception

CVPR 2025poster

Point cloud video understanding is becoming increasingly important in fields such as robotics, autonomous driving, and augmented reality, as they can accurately represent object motion and environmental changes. Despite the progress made in self-supervised learning methods for point cloud video unde…

2025

An Exploration with Entropy Constrained 3D Gaussians for 2D Video Compression

ICLR 2025poster

3D Gaussian Splatting (3DGS) has witnessed its rapid development in novel view synthesis, which attains high quality reconstruction and real-time rendering. At the same time, there is still a gap before implicit neural representation (INR) can become a practical compressor due to the lack of stream…

2025

AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing

CVPR 2025poster

Self-Supervised Video Hashing (SSVH) compresses videos into hash codes for efficient indexing and retrieval using unlabeled training videos. Existing approaches rely on random frame sampling to learn video features and treat all frames equally. This results in suboptimal hash codes, as it ignores fr…

2025

Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models

ACL 2025long

Large Audio-Language Models (LALMs), such as GPT-4o, have recently unlocked audio dialogue capabilities, enabling direct spoken exchanges with humans. The potential of LALMs broadens their applicability across a wide range of practical scenarios supported by audio dialogues. However, given these adv…

2025

CALF: Aligning LLMs for Time Series Forecasting via Cross-modal Fine-Tuning

AAAI 2025technical

Deep learning (e.g., Transformer) has been widely and successfully used in multivariate time series forecasting (MTSF). Unlike existing methods that focus on training models from a single modal of time series input, large language models (LLMs) based MTSF methods with cross-modal text and time serie…

2025

Cassic: Towards Content-Adaptive State-Space Models for Learned Image Compression

ICCV 2025poster

Learned image compression (LIC) demonstrates superior rate-distortion (RD) performance compared to traditional methods. Recent method MambaVC attempts to introduce Mamba, a variant of state space models, into this field aim to establish a new paradigm beyond convolutional neural networks and transfo…

Cited by 0SourcePDFScholar
2025

Diffusion Prior Interpolation for Flexibility Real-World Face Super-Resolution

AAAI 2025technical

Diffusion models represent the state-of-the-art in generative modeling. Due to their high training costs, many works leverage pre-trained diffusion models' powerful representations for downstream tasks, such as face super-resolution (FSR), through fine-tuning or prior-based methods. However, relying…

2025

DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation

ICLR 2025poster

Personalized image generation holds great promise in assisting humans in everyday work and life due to its impressive function in creatively generating personalized content. However, current evaluations either are automated but misalign with humans or require human evaluations that are time-consumin…

2025

Efficient Differentiable Approximation of Generalized Low-rank Regularization

IJCAI 2025

Low-rank regularization (LRR) has been widely applied in various machine learning tasks, but the associated optimization is challenging. Directly optimizing the rank function under constraints is NP-hard in general. To overcome this difficulty, various relaxations of the rank function were studied.

2025

Efficient Self-Supervised Video Hashing with Selective State Spaces

AAAI 2025technical

Self-supervised video hashing (SSVH) is a practical task in video indexing and retrieval. Although Transformers are predominant in SSVH for their impressive temporal modeling capabilities, they often suffer from computational and memory inefficiencies. Drawing inspiration from Mamba, an advanced sta…

2025

Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning

CVPR 2025poster

Visual In-Context Learning (VICL) enables adaptively solving vision tasks by leveraging pixel demonstrations, mimicking human-like task completion through analogy. Prompt selection is critical in VICL, but current methods assume the existence of a single "ideal" prompt in a pool of candidates, which…

2025

Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning

ICCV 2025poster

Partially Relevant Video Retrieval (PRVR) addresses the critical challenge of matching untrimmed videos with text queries describing only partial content. Existing methods suffer from geometric distortion in Euclidean space that sometimes misrepresents the intrinsic hierarchical structure of videos…

2025

Error-quantified Conformal Inference for Time Series

ICLR 2025poster

Uncertainty quantification in time series prediction is challenging due to the temporal dependence and distribution shift on sequential data. Conformal prediction provides a pivotal and flexible instrument for assessing the uncertainty of machine learning models through prediction sets. Recently, a…

2025

Expert-Enhanced Masked Point Modeling for Point Cloud Self-Supervised Learning

ICRA 2025

Recently, learning-based point cloud analysis has played a crucial role in robotic perception. Masked Point Modeling (MPM), owing to its powerful representational capabilities, has become the mainstream point cloud self-supervised learning method. However, existing MPM-based methods often suffer fro

Cited by 0SourcecodeScholar
2025

FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning

ICCV 2025poster

Visual Autoregressive (VAR) modeling has gained popularity for its shift towards next-scale prediction. However, existing VAR paradigms process the entire token map at each scale step, leading to the complexity and runtime scaling dramatically with image resolution. To address this challenge, we pro…

2025

GCD-Sampling: A General Cross-scale Decoupled Sampling for Point Cloud

AAAI 2025technical

Sampling strategy (e.g., fixed farthest point sampling) of point cloud has been an essential step for developing practical solutions in 3D computer vision tasks. Previous fixed sampling is simple, but suffer from suboptimal performance for downstream tasks. To adapt to target networks properly, adap…

2025

Going Beyond Feature Similarity: Effective Dataset distillation based on Class-aware Conditional Mutual Information

ICLR 2025poster

Dataset distillation (DD) aims to minimize the time and memory consumption needed for training deep neural networks on large datasets, by creating a smaller synthetic dataset that has similar performance to that of the full real dataset. However, current dataset distillation methods often result in…

2025

Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs

NeurIPS 2025poster

Large Vision-Language Models (LVLMs) are susceptible to hallucinations, where generated responses seem semantically plausible yet exhibit little or no relevance to the input image. Previous studies reveal that this issue primarily stems from LVLMs' over-reliance on language priors while disregarding…

Cited by 0SourceScholar
2025

Hierarchical Features Matter: A Deep Exploration of Progressive Parameterization Method for Dataset Distillation

CVPR 2025poster

Dataset distillation is an emerging dataset reduction method, which condenses large-scale datasets while maintaining task accuracy. Current parameterization methods achieve enhanced performance under extremely high compression ratio by optimizing determined synthetic dataset in informative feature d…

2025

IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion Models

ICML 2025poster

Fine-tuning pre-trained diffusion models under limited budgets has gained great success. In particular, the recent advances that directly fine-tune the quantized weights using Low-rank Adaptation (LoRA) further reduces training costs. Despite these progress, we point out that existing adaptation rec…

2025

LNeRV: Learnable Hierarchical Encoding Improve Neural Representation Video Codec

ICASSP 2025accepted

Existing Implicit Neural Representation (INR) video compression techniques have opened up new avenues in the field of video compression. NeRV maps the temporal coordinates to high-resolution images using neural networks, providing a more flexible and efficient encoding method for video data. However…

Cited by 1SourceScholar
2025

MambaIRv2: Attentive State Space Restoration

CVPR 2025poster

The Mamba-based image restoration backbones have recently demonstrated significant potential in balancing global reception and computational efficiency. However, the inherent causal modeling limitation of Mamba, where each token depends solely on its predecessors in the scanned sequence, restricts t…

2025

MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional Thresholds

EMNLP 2025

The rapid advancement of large language models has intensified public concerns about the potential misuse. Therefore, it is important to build trustworthy AI-generated text detection systems. Existing methods neglect stylistic modeling and mostly rely on static thresholds, which greatly limits the d

2025

Modeling Uncertainty in Composed Image Retrieval via Probabilistic Embeddings

ACL 2025long

Composed Image Retrieval (CIR) enables users to search for images using multimodal queries that combine text and reference images. While metric learning methods have shown promise, they rely on deterministic point embeddings that fail to capture the inherent uncertainty in the input data, in which u…

2025

One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training Models

ICCV 2025poster

Vision-Language Pre-training (VLP) models have exhibited unprecedented capability in many applications by taking full advantage of the learned multimodal alignment. However, previous studies have shown they are vulnerable to maliciously crafted adversarial samples. Despite recent success, these atta…

2025

PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba Adapter

CVPR 2025poster

Applying pre-trained models to assist point cloud understanding has recently become a mainstream paradigm in 3D perception. However, existing application strategies are straightforward, utilizing only the final output of the pre-trained model for various task heads. It neglects the rich complementar…

2025

Point Cloud Mixture-of-Domain-Experts Model for 3D Self-supervised Learning

IJCAI 2025

Point clouds, as a primary representation of 3D data, can be categorized into scene domain point clouds and object domain point clouds. Point cloud self-supervised learning (SSL) has become a mainstream paradigm for learning 3D representations. However, existing point cloud SSL primarily focuses on

Cited by 0SourcePDFScholar
2025

Pre-Trained Vision-Language Models as Noisy Partial Annotators

AAAI 2025technical

In noisy partial label learning, each training sample is associated with a set of candidate labels, and the ground-truth label may be contained within this set. With the emergence of powerful pre-trained vision-language models, e.g. CLIP, it is natural to consider using these models to automatically…

2025

Protecting Your Video Content: Disrupting Automated Video-based LLM Annotations

CVPR 2025poster

Recently, video-based large language models (video-based LLMs) have achieved impressive performance across various video comprehension tasks. However, this rapid advancement raises significant privacy and security concerns, particularly regarding the unauthorized use of personal video data in automa…

2025

RobNAS: Robust Neural Architecture Search for Point Cloud Adversarial Defense

ICASSP 2025accepted

As point clouds gain widespread application in fields such as autonomous driving and scene modeling, an increasing number of point cloud learning networks have emerged. As a result, research on 3D adversarial attacks and defenses has rapidly advanced. To the best of our knowledge, existing 3D defens…

Cited by 0SourceScholar
2025

Stealthy Shield Defense: A Conditional Mutual Information-Based Approach against Black-Box Model Inversion Attacks

ICLR 2025poster

Model inversion attacks (MIAs) aim to reconstruct the private training data by accessing the public model, raising concerns about privacy leakage. Black-box MIAs, where attackers can only query the model and obtain outputs, are closer to real-world scenarios. The latest black-box attacks have outper…

2025

TimeBridge: Non-Stationarity Matters for Long-term Time Series Forecasting

ICML 2025poster

Non-stationarity poses significant challenges for multivariate time series forecasting due to the inherent short-term fluctuations and long-term trends that can lead to spurious regressions or obscure essential long-term relationships. Most existing methods either eliminate or retain non-stationarit…

2025

TimeFilter: Patch-Specific Spatial-Temporal Graph Filtration for Time Series Forecasting

ICML 2025poster

Time series forecasting methods generally fall into two main categories: Channel Independent (CI) and Channel Dependent (CD) strategies. While CI overlooks important covariate relationships, CD captures all dependencies without distinction, introducing noise and reducing generalization. Recent advan…

2025

Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors

EMNLP 2025

The misuse of large language models (LLMs), such as academic plagiarism, has driven the development of detectors to identify LLM-generated texts. To bypass these detectors, paraphrase attacks have emerged to purposely rewrite these texts to evade detection. Despite the success, existing methods requ

2024

A Closer Look at GAN Priors: Exploiting Intermediate Features for Enhanced Model Inversion Attacks

ECCV 2024oral

"Model Inversion (MI) attacks aim to reconstruct privacy-sensitive training data from released models by utilizing output information, raising extensive concerns about the security of Deep Neural Networks (DNNs). Recent advances in generative adversarial networks (GANs) have contributed significantl…

2024

BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIP

CVPR 2024poster

Contrastive Vision-Language Pre-training known as CLIP has shown promising effectiveness in addressing downstream image recognition tasks. However recent works revealed that the CLIP model can be implanted with a downstream-oriented backdoor. On downstream tasks one victim model performs well on cle…

2024

BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional Bootstrapping

NeurIPS 2024poster

Adaptation of pretrained vision-language models such as CLIP to various downstream tasks have raised great interest in recent researches. Previous works have proposed a variety of test-time adaptation (TTA) methods to achieve strong generalization without any knowledge of the target domain. Howev…

2024

Boundary-aware Decoupled Flow Networks for Realistic Extreme Rescaling

IJCAI 2024poster

Recently developed generative methods, including invertible rescaling network (IRN) based and generative adversarial network (GAN) based methods, have demonstrated exceptional performance in image rescaling. However, IRN-based methods tend to produce over-smoothed results, while GAN-based methods ea…

2024

CAGEN: Controllable Anomaly Generator using Diffusion Model

ICASSP 2024accepted

Data augmentation has been widely applied in anomaly detection, which generates synthetic anomalous data for training. However, most existing anomaly augmentation methods focus on image-level cut-and-paste techniques, resulting in less realistic synthetic results, and are restricted to a few predefi…

Cited by 0SourceScholar
2024

CLIP-Guided Generative Networks for Transferable Targeted Adversarial Attacks

ECCV 2024poster

"Transferable targeted adversarial attacks aim to mislead models into outputting adversary-specified predictions in black-box scenarios. Recent studies have introduced single-target attacks that train a generator for each target class to generate highly transferable perturbations, resulting in subst…

2024

Controller-Guided Partial Label Consistency Regularization with Unlabeled Data

AAAI 2024technical

Partial label learning (PLL) learns from training examples each associated with multiple candidate labels, among which only one is valid. In recent years, benefiting from the strong capability of dealing with ambiguous supervision and the impetus of modern data augmentation methods, consistency regu…

Cited by 3SourcePDFScholar
2024

DDN: Dual-domain Dynamic Normalization for Non-stationary Time Series Forecasting

NeurIPS 2024poster

Deep neural networks (DNNs) have recently achieved remarkable advancements in time series forecasting (TSF) due to their powerful ability of sequence dependence modeling. To date, existing DNN-based TSF methods still suffer from unreliable predictions for real-world data due to its non-stationarity…

Cited by 3SourcePDFScholar
2024

Everyday Object Meets Vision-and-Language Navigation Agent via Backdoor

NeurIPS 2024poster

Vision-and-Language Navigation (VLN) requires an agent to dynamically explore environments following natural language. The VLN agent, closely integrated into daily lives, poses a substantial threat to the security of privacy and property upon the occurrence of malicious behavior. However, this serio…

Cited by 0SourcePDFScholar
2024

GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video Retrieval

AAAI 2024technical

Given a text query, partially relevant video retrieval (PRVR) seeks to find untrimmed videos containing pertinent moments in a database. For PRVR, clip modeling is essential to capture the partial relationship between texts and videos. Current PRVR methods adopt scanning-based clip construction to a…

2024

GladCoder: Stylized QR Code Generation with Grayscale-Aware Denoising Process

IJCAI 2024poster

Traditional QR codes consist of a grid of black-and-white square modules, which lack aesthetic appeal and meaning for human perception. This has motivated recent research to beautify the visual appearance of QR codes. However, there exists a trade-off between the visual quality and scanning-robustne…

Cited by 0SourcePDFScholar
2024

Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

ICLR 2024poster

Large vision-language models (VLMs) such as GPT-4 have achieved exceptional performance across various multi-modal tasks. However, the deployment of VLMs necessitates substantial energy consumption and computational resources. Once attackers maliciously induce high energy consumption and latency tim…

2024

LCM: Locally Constrained Compact Point Cloud Model for Masked Point Modeling

NeurIPS 2024poster

The pre-trained point cloud model based on Masked Point Modeling (MPM) has exhibited substantial improvements across various tasks. However, these models heavily rely on the Transformer, leading to quadratic complexity and limited decoder, hindering their practice application. To address this limita…

2024

MambaIR: A Simple Baseline for Image Restoration with State-Space Model

ECCV 2024poster

"Recent years have seen significant advancements in image restoration, largely attributed to the development of modern deep neural networks, such as CNNs and Transformers. However, existing restoration backbones often face the dilemma between global receptive fields and efficient computation, hinder…

2024

Not All Prompts Are Secure: A Switchable Backdoor Attack Against Pre-trained Vision Transfomers

CVPR 2024poster

Given the power of vision transformers a new learning paradigm pre-training and then prompting makes it more efficient and effective to address downstream visual recognition tasks. In this paper we identify a novel security threat towards such a paradigm from the perspective of backdoor attacks. Spe…

2024

Parameter Efficient Adaptation for Image Restoration with Heterogeneous Mixture-of-Experts

NeurIPS 2024poster

Designing single-task image restoration models for specific degradation has seen great success in recent years. To achieve generalized image restoration, all-in-one methods have recently been proposed and shown potential for multiple restoration tasks using one single model. Despite the promising re…

2024

Parameter-Efficient and Memory-Efficient Tuning for Vision Transformer: A Disentangled Approach

ECCV 2024poster

"Recent works on parameter-efficient transfer learning (PETL) show the potential to adapt a pre-trained Vision Transformer to downstream recognition tasks with only a few learnable parameters. However, since they usually insert new structures into the pre-trained model, entire intermediate features…

2024

Periodicity Decoupling Framework for Long-term Series Forecasting

ICLR 2024poster

Convolutional neural network (CNN)-based and Transformer-based methods have recently made significant strides in time series forecasting, which excel at modeling local temporal variations or capturing long-term dependencies. However, real-world time series usually contain intricate temporal patterns…

2024

ReFIR: Grounding Large Restoration Models with Retrieval Augmentation

NeurIPS 2024poster

Recent advances in diffusion-based Large Restoration Models (LRMs) have significantly improved photo-realistic image restoration by leveraging the internal knowledge embedded within model weights. However, existing LRMs often suffer from the hallucination dilemma, i.e., producing incorrect contents…

2024

Towards Compact 3D Representations via Point Feature Enhancement Masked Autoencoders

AAAI 2024technical

Learning 3D representation plays a critical role in masked autoencoder (MAE) based pre-training methods for point cloud, including single-modal and cross-modal based MAE. Specifically, although cross-modal MAE methods learn strong 3D representations via the auxiliary of other modal knowledge, they…

2024

Towards Faithful XAI Evaluation via Generalization-Limited Backdoor Watermark

ICLR 2024poster

Saliency-based representation visualization (SRV) ($e.g.$, Grad-CAM) is one of the most classical and widely adopted explainable artificial intelligence (XAI) methods for its simplicity and efficiency. It can be used to interpret deep neural networks by locating saliency areas contributing the most…

2024

Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding

AAAI 2024technical

In recent years, vision language pre-training frameworks have made significant progress in natural language processing and computer vision, achieving remarkable performance improvement on various downstream tasks. However, when extended to point cloud data, existing works mainly focus on building ta…

2024

WFTNet: Exploiting Global and Local Periodicity in Long-Term Time Series Forecasting

ICASSP 2024accepted

Recent CNN and Transformer-based models tried to utilize frequency and periodicity information for long-term time series forecasting. However, most existing work is based on Fourier transform, which cannot capture fine-grained and local frequency structure. In this paper, we propose a Wavelet-Fourie…

Cited by 0SourceScholar
2023

Backdoor Defense via Adaptively Splitting Poisoned Dataset

CVPR 2023poster

Backdoor defenses have been studied to alleviate the threat of deep neural networks (DNNs) being backdoor attacked and thus maliciously altered. Since DNNs usually adopt some external training data from an untrusted third party, a robust backdoor defense strategy during the training stage is of impo…

2023

Combating Unknown Bias with Effective Bias-Conflicting Scoring and Gradient Alignment

AAAI 2023technical

Models notoriously suffer from dataset biases which are detrimental to robustness and generalization. The identify-emphasize paradigm shows a promising effect in dealing with unknown biases. However, we find that it is still plagued by two challenges: A, the quality of the identified bias-conflictin…

Cited by 9SourcePDFScholar
2023

Contrastive Masked Autoencoders for Self-Supervised Video Hashing

AAAI 2023technical

Self-Supervised Video Hashing (SSVH) models learn to generate short binary representations for videos without ground-truth supervision, facilitating large-scale video retrieval efficiency and attracting increasing research attention. The success of SSVH lies in the understanding of video content and…

2023

Difficulty-Aware Data Augmentor for Scene Text Recognition

ICASSP 2023accepted

Deep neural network (DNN) based scene text recognition (STR) methods usually require a large amount of annotated data for training, which is time-consuming and cost-expensive in practice. To address this issue, many data augmentation methods have been developed to train recognizers by improving the…

Cited by 0SourceScholar
2023

Domain Watermark: Effective and Harmless Dataset Copyright Protection is Closed at Hand

NeurIPS 2023poster

The prosperity of deep neural networks (DNNs) is largely benefited from open-source datasets, based on which users can evaluate and improve their methods. In this paper, we revisit backdoor-based dataset ownership verification (DOV), which is currently the only feasible approach to protect the copyr…

2023

FSR: A General Frequency-Oriented Framework to Accelerate Image Super-resolution Networks

AAAI 2023technical

Deep neural networks (DNNs) have witnessed remarkable achievement in image super-resolution (SR), and plenty of DNN-based SR models with elaborated network designs have recently been proposed. However, existing methods usually require substantial computations by operating in spatial domain. To addre…

2023

GIFD: A Generative Gradient Inversion Method with Feature Domain Optimization

ICCV 2023poster

Federated Learning (FL) has recently emerged as a promising distributed machine learning framework to preserve clients' privacy, by allowing multiple clients to upload the gradients calculated from their local data to a central server. Recent studies find that the exchanged gradients also take the r…

Cited by 40PDFcodeScholar
2023

Instance-aware Dynamic Prompt Tuning for Pre-trained Point Cloud Models

ICCV 2023poster

Pre-trained point cloud models have found extensive applications in 3D understanding tasks like object classification and part segmentation. However, the prevailing strategy of full fine-tuning in downstream tasks leads to large per-task storage overhead for model parameters, which limits the effici…

Cited by 48PDFcodeScholar
2023

Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection

ICLR 2023top-5%

The exploration problem is one of the main challenges in deep reinforcement learning (RL). Recent promising works tried to handle the problem with population-based methods, which collect samples with diverse behaviors derived from a population of different exploratory policies. Adaptive policy selec…

Cited by 17SourcePDFScholar
2023

Learned Distributed Image Compression with Multi-Scale Patch Matching in Feature Domain

AAAI 2023technical

Beyond achieving higher compression efficiency over classical image compression codecs, deep image compression is expected to be improved with additional side information, e.g., another image from a different perspective of the same scene. To better utilize the side information under the distributed…

Cited by 13SourcePDFScholar
2023

Learning Transferable Spatiotemporal Representations From Natural Script Knowledge

CVPR 2023poster

Pre-training on large-scale video data has become a common recipe for learning transferable spatiotemporal representations in recent years. Despite some progress, existing methods are mostly limited to highly curated datasets (e.g., K400) and exhibit unsatisfactory out-of-the-box representations. We…

2023

One-bit Flip is All You Need: When Bit-flip Attack Meets Model Training

ICCV 2023poster

Deep neural networks (DNNs) are widely deployed on real-world devices. Concerns regarding their security have gained great attention from researchers. Recently, a new weight modification attack called bit flip attack (BFA) was proposed, which exploits memory fault inject techniques such as row hamme…

Cited by 19PDFcodeScholar
2023

Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

AAAI 2023technical

Real-world recognition system often encounters the challenge of unseen labels. To identify such unseen labels, multi-label zero-shot learning (ML-ZSL) focuses on transferring knowledge by a pre-trained textual label embedding (e.g., GloVe). However, such methods only exploit single-modal knowledge f…

2023

SFR: Semantic-Aware Feature Rendering of Point Cloud

ICASSP 2023accepted

Multi-view projection methods have demonstrated their ability to reach state-of-the-art performance in point cloud downstream tasks(e.g., classification and retrieval). These methods first require rendering the point cloud into 2D multi-view images. However, conventional methods only project the geo…

Cited by 0SourceScholar
2023

Semantic Preserving Learning for Task-Oriented Point Cloud Downsampling

ICASSP 2023accepted

Recent years have witnessed a tremendous growth in the scale and resolution of point clouds. To facilitate the applications of point cloud in downsampling tasks (e.g., point cloud classification), several task-oriented downsampling works have been developed by training with the task-specific loss wi…

Cited by 0SourceScholar
2023

Towards Robust Model Watermark via Reducing Parametric Vulnerability

ICCV 2023poster

Deep neural networks are valuable assets considering their commercial benefits and huge demands for costly annotation and computation resources. To protect the copyright of DNNs, backdoor-based ownership verification becomes popular recently, in which the model owner can watermark the model by embed…

Cited by 11PDFcodeScholar
2023

Towards Robust Scene Text Image Super-resolution via Explicit Location Enhancement

IJCAI 2023poster

Scene text image super-resolution (STISR), aiming to improve image quality while boosting downstream scene text recognition accuracy, has recently achieved great success. However, most existing methods treat the foreground (character regions) and background (non-character regions) equally in the for…

2023

Unsupervised Surface Anomaly Detection with Diffusion Probabilistic Model

ICCV 2023poster

Unsupervised surface anomaly detection aims at discovering and localizing anomalous patterns using only anomaly-free training samples. Reconstruction-based models are among the most popular and successful methods, which rely on the assumption that anomaly regions are more difficult to reconstruct. H…

Cited by 77PDFScholar
2022

Boosting Black-Box Attack With Partially Transferred Conditional Adversarial Distribution

CVPR 2022poster

This work studies black-box adversarial attacks against deep neural networks (DNNs), where the attacker can only access the query feedback returned by the attacked DNN model, while other information such as model parameters or the training datasets are unknown. One promising approach to improve atta…

Cited by 49PDFcodeScholar
2022

Contrastive Quantization with Code Memory for Unsupervised Image Retrieval

AAAI 2022technical

The high efficiency in computation and storage makes hashing (including binary hashing and quantization) a common strategy in large-scale retrieval systems. To alleviate the reliance on expensive annotations, unsupervised deep hashing becomes an important research problem. This paper provides a nove…

2022

Defending against Model Stealing via Verifying Embedded External Features

AAAI 2022technical

Obtaining a well-trained model involves expensive data collection and training procedures, therefore the model is a valuable intellectual property. Recent studies revealed that adversaries can `steal' deployed models even when they have no training samples and can not get access to the model paramet…

2022

Few-Shot Backdoor Attacks on Visual Object Tracking

ICLR 2022poster

Visual object tracking (VOT) has been widely adopted in mission-critical applications, such as autonomous driving and intelligent surveillance systems. In current practice, third-party resources such as datasets, backbone networks, and training platforms are frequently used to train high-performance…

2022

Hardly Perceptible Trojan Attack against Neural Networks with Bit Flips

ECCV 2022poster

"The security of deep neural networks (DNNs) has attracted increasing attention due to their widespread use in various applications. Recently, the deployed DNNs have been demonstrated to be vulnerable to Trojan attacks, which manipulate model parameters with bit flips to inject a hidden behavior and…

2022

Improving Vision Transformers by Revisiting High-Frequency Components

ECCV 2022poster

"The transformer models have shown promising effectiveness in dealing with various vision tasks. However, compared with training Convolutional Neural Network (CNN) models, training Vision Transformer (ViT) models is more difficult and relies on the large-scale training set. To explain this observati…

2022

NeXT: Towards High Quality Neural Radiance Fields via Multi-Skip Transformer

ECCV 2022poster

"Neural Radiance Fields (NeRF) methods show impressive performance for novel view synthesis by representing a scene via a neural network. However, most existing NeRF based methods, including its variants, treat each sample point individually as input, while ignoring the inherent relationships betwee…

2022

PILC: Practical Image Lossless Compression With an End-to-End GPU Oriented Neural Framework

CVPR 2022poster

Generative model based image lossless compression algorithms have seen a great success in improving compression ratio. However, the throughput for most of them is less than 1 MB/s even with the most advanced AI accelerated chips, preventing them from most real-world applications, which often require…

Cited by 25PDFScholar
2022

SimCC: A Simple Coordinate Classification Perspective for Human Pose Estimation

ECCV 2022poster

"The 2D heatmap-based approaches have dominated Human Pose Estimation (HPE) for years due to high performance. However, the long-standing quantization error problem in the 2D heatmap-based methods leads to several well-known drawbacks: 1) The performance for the low-resolution inputs is limited; 2)…

2022

Untargeted Backdoor Watermark: Towards Harmless and Stealthy Dataset Copyright Protection

NeurIPS 2022accept

Deep neural networks (DNNs) have demonstrated their superiority in practice. Arguably, the rapid development of DNNs is largely benefited from high-quality (open-sourced) datasets, based on which researchers and developers can easily evaluate and improve their learning methods. Since the data collec…

2021

Attention on Attention Sparse Dense Convolutional Network for Financial Signal Processing

ICASSP 2021accepted

Financial signal processing is a matter of great concern in FinTech. Traditionally, recurrent networks are often used to model time series, while the latest research shows that convolutional networks, especially temporal convolutional networks (TCNs), are also powerful and effective for a large numb…

Cited by 0SourceScholar
2021

Backdoor Attack Against Speaker Verification

ICASSP 2021accepted

Speaker verification has been widely and successfully adopted in many mission-critical areas for user identification. The training of speaker verification requires a large amount of data, therefore users usually need to adopt third-party data (e.g., data from the Internet or third-party data company…

Cited by 0SourceScholar
2021

Clustering Effect of Adversarial Robust Models

NeurIPS 2021spotlight

Adversarial robustness has received increasing attention along with the study of adversarial examples. So far, existing works show that robust models not only obtain robustness against various adversarial attacks but also boost the performance in some downstream tasks. However, the underlying mechan…

2021

H-GPR: A Hybrid Strategy for Large-Scale Gaussian Process Regression

ICASSP 2021accepted

With the massive volume of data emerging from both scientific and industrial domains, it has become a desideratum to improve the scalability of Gaussian process regression (GPR). There are two major approaches to assuage its $\mathcal{O}\left( {{n^3}} \right)$ training complexity: the aggregation ba…

Cited by 0SourceScholar
2021

HOCA: Higher-Order Channel Attention for Single Image Super-Resolution

ICASSP 2021accepted

Convolutional neural networks (CNNs) have obtained great success in single image super-resolution (SR). More recent works (e.g., RCAN and SAN) have obtained remarkable performance with channel attention based on first- or second-order statistics of features. However, these methods neglect the rich f…

Cited by 5SourceScholar
2021

Improving Adversarial Robustness via Channel-wise Activation Suppressing

ICLR 2021spotlight

The study of adversarial examples and their activations have attracted significant attention for secure and robust learning with deep neural networks (DNNs). Different from existing works, in this paper, we highlight two new characteristics of adversarial examples from the channel-wise activation p…

2021

Targeted Attack against Deep Neural Networks via Flipping Limited Weight Bits

ICLR 2021poster

To explore the vulnerability of deep neural networks (DNNs), many attack paradigms have been well studied, such as the poisoning-based backdoor attack in the training stage and the adversarial attack in the inference stage. In this paper, we study a novel attack paradigm, which modifies model parame…

2021

TokenPose: Learning Keypoint Tokens for Human Pose Estimation

ICCV 2021poster

Human pose estimation deeply relies on visual clues and anatomical constraints between parts to locate keypoints. Most existing CNN-based methods do well in visual representation, however, lacking in the ability to explicitly learn the constraint relationships between keypoints. In this paper, we pr…

Cited by 386PDFcodeScholar
2021

t-k-means: A ROBUST AND STABLE k-means VARIANT

ICASSP 2021accepted

k-means algorithm is one of the most classical clustering methods, which has been widely and successfully used in signal processing. However, due to the thin-tailed property of the Gaussian distribution, k-means algorithm suffers from relatively poor performance on the dataset containing heavy-taile…

Cited by 0SourceScholar
2020

Hijacking Tracker: A Powerful Adversarial Attack on Visual Tracking

ICASSP 2020accepted

Visual object tracking has made important breakthroughs with the assistance of deep learning models. Unfortunately, recent research has clearly proved that deep learning models are vulnerable to malicious adversarial attacks, which mislead the models making wrong decisions by perturbing the input im…

Cited by 0SourceScholar
2020

Improving Query Efficiency of Black-box Adversarial Attack

ECCV 2020poster

Deep neural networks (DNNs) have demonstrated excellent performance on various tasks, however they are under the risk of adversarial examples that can be easily generated when the target model is accessible to an attacker (white-box setting). As plenty of machine learning models have been deployed v…

2020

Maintaining Discrimination and Fairness in Class Incremental Learning

CVPR 2020poster

Deep neural networks (DNNs) have been applied in class incremental learning, which aims to solve common real-world problems of learning new classes continually. One drawback of standard DNNs is that they are prone to catastrophic forgetting. Knowledge distillation (KD) is a commonly used technique t…

Cited by 614PDFScholar
2020

One-Shot Adversarial Attacks on Visual Tracking With Dual Attention

CVPR 2020poster

Almost all adversarial attacks in computer vision are aimed at pre-known object categories, which could be offline trained for generating perturbations. But as for visual object tracking, the tracked target categories are normally unknown in advance. However, the tracking algorithms also have potent…

Cited by 100PDFScholar
2020

Skip Connections Matter: On the Transferability of Adversarial Examples Generated with ResNets

ICLR 2020spotlight

Skip connections are an essential component of current state-of-the-art deep neural networks (DNNs) such as ResNet, WideResNet, DenseNet, and ResNeXt. Despite their huge success in building deeper and more powerful DNNs, we identify a surprising \emph{security weakness} of skip connections in this p…

Cited by 423SourcecodeScholar
2020

Stochastic Deep Gaussian Processes over Graphs

NeurIPS 2020poster

In this paper we propose Stochastic Deep Gaussian Processes over Graphs (DGPG), which are deep structure models that learn the mappings between input and output signals in graph domains. The approximate posterior distributions of the latent variables are derived with variational inference, and the e…

2020

Targeted Attack for Deep Hashing based Retrieval

ECCV 2020poster

The deep hashing based retrieval method is widely adopted in large-scale image and video retrieval. However, there is little investigation on its security. In this paper, we propose a novel method, dubbed deep hashing targeted attack (DHTA), to study the targeted attack on such retrieval. Specifical…

2020

Training Interpretable Convolutional Neural Networks by Differentiating Class-specific Filters

ECCV 2020poster

Convolutional neural networks (CNNs) have been successfully used in a range of tasks. However, CNNs are often viewed as ""black-box"" and lack of interpretability. One main reason is due to the filter-class entanglement -- an intricate many-to-many correspondence between filters and classes. Most ex…

2019

Hilbert-Based Generative Defense for Adversarial Examples

ICCV 2019poster

Adversarial perturbations of clean images are usually imperceptible for human eyes, but can confidently fool deep neural networks (DNNs) to make incorrect predictions. Such vulnerability of DNNs raises serious security concerns about their practicability in security-sensitive applications. To defend…

Cited by 62PDFScholar
2019

Second-Order Attention Network for Single Image Super-Resolution

CVPR 2019oral

Recently, deep convolutional neural networks (CNNs) have been widely explored in single image super-resolution (SISR) and obtained remarkable performance. However, most of the existing CNN-based SISR methods mainly focus on wider or deeper architecture design, neglecting to explore the feature corre…

Cited by 2042PDFScholar
2018

BML: A High-performance, Low-cost Gradient Synchronization Algorithm for DML Training

NeurIPS 2018poster

In distributed machine learning (DML), the network performance between machines significantly impacts the speed of iterative training. In this paper we propose BML, a new gradient synchronization algorithm with higher network performance and lower network cost than the current practice. BML runs on…

Cited by 39SourcePDFScholar
2018

Iterative Learning With Open-Set Noisy Labels

CVPR 2018poster

Large-scale datasets possessing clean label annotations are crucial for training Convolutional Neural Networks (CNNs). However, labeling large-scale data can be very costly and error-prone, and even high-quality datasets are likely to contain noisy (incorrect) labels. Existing works usually employ a…

2017

Accelerated Stochastic Greedy Coordinate Descent by Soft Thresholding Projection onto Simplex

NeurIPS 2017spotlight

In this paper we study the well-known greedy coordinate descent (GCD) algorithm to solve $\ell_1$-regularized problems and improve GCD by the two popular strategies: Nesterov's acceleration and stochastic optimization. Firstly, we propose a new rule for greedy selection based on an $\ell_1$-norm sq…

Cited by 16SourcePDFScholar