← Search

Tao Dai

68 accepted papers

2026

CASL: Curvature-Augmented Self-supervised Learning for 3D Anomaly Detection

AAAI 2026technical

Deep learning-based 3D anomaly detection methods have demonstrated significant potential in industrial manufacturing. However, many approaches are specifically designed for anomaly detection tasks, which limits their generalizability to other 3D tasks. In contrast, self-supervised point cloud models

Cited by 0SourcePDFScholar
2026

Improving Deepfake Detection with Reinforcement Learning-Based Adaptive Data Augmentation

AAAI 2026technical

The generalization capability of deepfake detectors is crucial for real-world applications. Data augmentation to generate synthetic fake faces has served as an effective strategy to enhance generalization. Interestingly, current state-of-the-art (SoTA) methods rely on fixed augmentation strategies,

Cited by 0SourcePDFScholar
2026

RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization

ICLR 2026poster

Visual manipulation localization (VML) aims to identify tampered regions in images and videos, a task that has become increasingly challenging with the rise of advanced editing tools. Existing methods face two main issues: resolution diversity, where resizing or padding distorts forensic traces and…

Cited by 0SourcecodeScholar
2026

SparseEval: Efficient Evaluation of Large Language Models by Sparse Optimization

ICLR 2026poster

As large language models (LLMs) continue to scale up, their performance on various downstream tasks has significantly improved. However, evaluating their capabilities has become increasingly expensive, as performing inference on a large number of benchmark samples incurs high computational costs. In…

Cited by 0SourcecodeScholar
2025

3D-LMVIC: Learning-based Multi-View Image Compression with 3D Gaussian Geometric Priors

ICML 2025poster

Existing multi-view image compression methods often rely on 2D projection-based similarities between views to estimate disparities. While effective for small disparities, such as those in stereo images, these methods struggle with the more complex disparities encountered in wide-baseline multi-camer…

Cited by 0SourcePDFScholar
2025

Adapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal Perception

CVPR 2025poster

Point cloud video understanding is becoming increasingly important in fields such as robotics, autonomous driving, and augmented reality, as they can accurately represent object motion and environmental changes. Despite the progress made in self-supervised learning methods for point cloud video unde…

2025

Adaptive Multi-Scale Decomposition Framework for Time Series Forecasting

AAAI 2025technical

Transformer-based and MLP-based methods have emerged as leading approaches in time series forecasting (TSF). However, real-world time series often show different patterns at different scales, and future changes are shaped by the interplay of these overlapping scales, requiring high-capacity models.…

2025

CALF: Aligning LLMs for Time Series Forecasting via Cross-modal Fine-Tuning

AAAI 2025technical

Deep learning (e.g., Transformer) has been widely and successfully used in multivariate time series forecasting (MTSF). Unlike existing methods that focus on training models from a single modal of time series input, large language models (LLMs) based MTSF methods with cross-modal text and time serie…

2025

Cassic: Towards Content-Adaptive State-Space Models for Learned Image Compression

ICCV 2025poster

Learned image compression (LIC) demonstrates superior rate-distortion (RD) performance compared to traditional methods. Recent method MambaVC attempts to introduce Mamba, a variant of state space models, into this field aim to establish a new paradigm beyond convolutional neural networks and transfo…

Cited by 0SourcePDFScholar
2025

DCSF-KD: Dynamic Channel-wise Spatial Feature Knowledge Distillation for Object Detection

AAAI 2025technical

Knowledge distillation (KD) has recently gained great success in the field of object detection. By transferring the knowledge of the spatial or channel domain from the teacher model to the student model, it allows for a more compact representation with minimal performance loss. Despite this progress…

2025

DIIN: Diffusion Iterative Implicit Networks for Arbitrary-scale Super-resolution

IJCAI 2025

Implicit neural representation (INR) aims to represent continuous domain signals via implicit neural functions and has achieved great success in arbitrary-scale image super-resolution (SR). However, most existing INR-based SR methods focus on learning implicit features from independent coordinate, w

2025

Diffusion Prior Interpolation for Flexibility Real-World Face Super-Resolution

AAAI 2025technical

Diffusion models represent the state-of-the-art in generative modeling. Due to their high training costs, many works leverage pre-trained diffusion models' powerful representations for downstream tasks, such as face super-resolution (FSR), through fine-tuning or prior-based methods. However, relying…

2025

Efficient Differentiable Approximation of Generalized Low-rank Regularization

IJCAI 2025

Low-rank regularization (LRR) has been widely applied in various machine learning tasks, but the associated optimization is challenging. Directly optimizing the rank function under constraints is NP-hard in general. To overcome this difficulty, various relaxations of the rank function were studied.

2025

Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning

CVPR 2025poster

Visual In-Context Learning (VICL) enables adaptively solving vision tasks by leveraging pixel demonstrations, mimicking human-like task completion through analogy. Prompt selection is critical in VICL, but current methods assume the existence of a single "ideal" prompt in a pool of candidates, which…

2025

EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models

AAAI 2025technical

Vision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries to retrieve corresponding images (or text). However, existing methods often neglect the rich visual semantics knowledge of entities, thus leading to inc…

Cited by 0SourcePDFScholar
2025

Expert-Enhanced Masked Point Modeling for Point Cloud Self-Supervised Learning

ICRA 2025

Recently, learning-based point cloud analysis has played a crucial role in robotic perception. Masked Point Modeling (MPM), owing to its powerful representational capabilities, has become the mainstream point cloud self-supervised learning method. However, existing MPM-based methods often suffer fro

Cited by 0SourcecodeScholar
2025

FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning

ICCV 2025poster

Visual Autoregressive (VAR) modeling has gained popularity for its shift towards next-scale prediction. However, existing VAR paradigms process the entire token map at each scale step, leading to the complexity and runtime scaling dramatically with image resolution. To address this challenge, we pro…

2025

GCD-Sampling: A General Cross-scale Decoupled Sampling for Point Cloud

AAAI 2025technical

Sampling strategy (e.g., fixed farthest point sampling) of point cloud has been an essential step for developing practical solutions in 3D computer vision tasks. Previous fixed sampling is simple, but suffer from suboptimal performance for downstream tasks. To adapt to target networks properly, adap…

2025

IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion Models

ICML 2025poster

Fine-tuning pre-trained diffusion models under limited budgets has gained great success. In particular, the recent advances that directly fine-tune the quantized weights using Low-rank Adaptation (LoRA) further reduces training costs. Despite these progress, we point out that existing adaptation rec…

2025

LNeRV: Learnable Hierarchical Encoding Improve Neural Representation Video Codec

ICASSP 2025accepted

Existing Implicit Neural Representation (INR) video compression techniques have opened up new avenues in the field of video compression. NeRV maps the temporal coordinates to high-resolution images using neural networks, providing a more flexible and efficient encoding method for video data. However…

Cited by 1SourceScholar
2025

MC3D-AD: A Unified Geometry-aware Reconstruction Model for Multi-category 3D Anomaly Detection

IJCAI 2025

3D Anomaly Detection (AD) is a promising means of controlling the quality of manufactured products. However, existing methods typically require carefully training a task-specific model for each category independently, leading to high cost, low efficiency, and weak generalization. This study presents

2025

MambaIRv2: Attentive State Space Restoration

CVPR 2025poster

The Mamba-based image restoration backbones have recently demonstrated significant potential in balancing global reception and computational efficiency. However, the inherent causal modeling limitation of Mamba, where each token depends solely on its predecessors in the scanned sequence, restricts t…

2025

PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba Adapter

CVPR 2025poster

Applying pre-trained models to assist point cloud understanding has recently become a mainstream paradigm in 3D perception. However, existing application strategies are straightforward, utilizing only the final output of the pre-trained model for various task heads. It neglects the rich complementar…

2025

Point Cloud Mixture-of-Domain-Experts Model for 3D Self-supervised Learning

IJCAI 2025

Point clouds, as a primary representation of 3D data, can be categorized into scene domain point clouds and object domain point clouds. Point cloud self-supervised learning (SSL) has become a mainstream paradigm for learning 3D representations. However, existing point cloud SSL primarily focuses on

Cited by 0SourcePDFScholar
2025

Protecting Your Video Content: Disrupting Automated Video-based LLM Annotations

CVPR 2025poster

Recently, video-based large language models (video-based LLMs) have achieved impressive performance across various video comprehension tasks. However, this rapid advancement raises significant privacy and security concerns, particularly regarding the unauthorized use of personal video data in automa…

2025

TimeBridge: Non-Stationarity Matters for Long-term Time Series Forecasting

ICML 2025poster

Non-stationarity poses significant challenges for multivariate time series forecasting due to the inherent short-term fluctuations and long-term trends that can lead to spurious regressions or obscure essential long-term relationships. Most existing methods either eliminate or retain non-stationarit…

2025

TimeFilter: Patch-Specific Spatial-Temporal Graph Filtration for Time Series Forecasting

ICML 2025poster

Time series forecasting methods generally fall into two main categories: Channel Independent (CI) and Channel Dependent (CD) strategies. While CI overlooks important covariate relationships, CD captures all dependencies without distinction, introducing noise and reducing generalization. Recent advan…

2024

BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional Bootstrapping

NeurIPS 2024poster

Adaptation of pretrained vision-language models such as CLIP to various downstream tasks have raised great interest in recent researches. Previous works have proposed a variety of test-time adaptation (TTA) methods to achieve strong generalization without any knowledge of the target domain. Howev…

2024

Boundary-aware Decoupled Flow Networks for Realistic Extreme Rescaling

IJCAI 2024poster

Recently developed generative methods, including invertible rescaling network (IRN) based and generative adversarial network (GAN) based methods, have demonstrated exceptional performance in image rescaling. However, IRN-based methods tend to produce over-smoothed results, while GAN-based methods ea…

2024

CLIP-Guided Generative Networks for Transferable Targeted Adversarial Attacks

ECCV 2024poster

"Transferable targeted adversarial attacks aim to mislead models into outputting adversary-specified predictions in black-box scenarios. Recent studies have introduced single-target attacks that train a generator for each target class to generate highly transferable perturbations, resulting in subst…

2024

DDN: Dual-domain Dynamic Normalization for Non-stationary Time Series Forecasting

NeurIPS 2024poster

Deep neural networks (DNNs) have recently achieved remarkable advancements in time series forecasting (TSF) due to their powerful ability of sequence dependence modeling. To date, existing DNN-based TSF methods still suffer from unreliable predictions for real-world data due to its non-stationarity…

Cited by 3SourcePDFScholar
2024

FreqFormer: Frequency-aware Transformer for Lightweight Image Super-resolution

IJCAI 2024poster

Transformer-based models have been widely and successfully used in various low-vision visual tasks, and have achieved remarkable performance in single image super-resolution (SR). Despite the significant progress in SR, Transformer-based SR methods (e.g., SwinIR) still suffer from the problems of…

2024

GladCoder: Stylized QR Code Generation with Grayscale-Aware Denoising Process

IJCAI 2024poster

Traditional QR codes consist of a grid of black-and-white square modules, which lack aesthetic appeal and meaning for human perception. This has motivated recent research to beautify the visual appearance of QR codes. However, there exists a trade-off between the visual quality and scanning-robustne…

Cited by 0SourcePDFScholar
2024

LCM: Locally Constrained Compact Point Cloud Model for Masked Point Modeling

NeurIPS 2024poster

The pre-trained point cloud model based on Masked Point Modeling (MPM) has exhibited substantial improvements across various tasks. However, these models heavily rely on the Transformer, leading to quadratic complexity and limited decoder, hindering their practice application. To address this limita…

2024

Parameter Efficient Adaptation for Image Restoration with Heterogeneous Mixture-of-Experts

NeurIPS 2024poster

Designing single-task image restoration models for specific degradation has seen great success in recent years. To achieve generalized image restoration, all-in-one methods have recently been proposed and shown potential for multiple restoration tasks using one single model. Despite the promising re…

2024

Periodicity Decoupling Framework for Long-term Series Forecasting

ICLR 2024poster

Convolutional neural network (CNN)-based and Transformer-based methods have recently made significant strides in time series forecasting, which excel at modeling local temporal variations or capturing long-term dependencies. However, real-world time series usually contain intricate temporal patterns…

2024

Procedural Level Generation with Diffusion Models from a Single Example

AAAI 2024technical

Level generation is a central focus of Procedural Content Generation (PCG), yet deep learning-based approaches are limited by scarce training data, i.e., human-designed levels. Despite being a dominant framework, Generative Adversarial Networks (GANs) exhibit a substantial quality gap between genera…

2024

ReFIR: Grounding Large Restoration Models with Retrieval Augmentation

NeurIPS 2024poster

Recent advances in diffusion-based Large Restoration Models (LRMs) have significantly improved photo-realistic image restoration by leveraging the internal knowledge embedded within model weights. However, existing LRMs often suffer from the hallucination dilemma, i.e., producing incorrect contents…

2024

Towards Compact 3D Representations via Point Feature Enhancement Masked Autoencoders

AAAI 2024technical

Learning 3D representation plays a critical role in masked autoencoder (MAE) based pre-training methods for point cloud, including single-modal and cross-modal based MAE. Specifically, although cross-modal MAE methods learn strong 3D representations via the auxiliary of other modal knowledge, they…

2024

Towards Faithful XAI Evaluation via Generalization-Limited Backdoor Watermark

ICLR 2024poster

Saliency-based representation visualization (SRV) ($e.g.$, Grad-CAM) is one of the most classical and widely adopted explainable artificial intelligence (XAI) methods for its simplicity and efficiency. It can be used to interpret deep neural networks by locating saliency areas contributing the most…

2024

Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding

AAAI 2024technical

In recent years, vision language pre-training frameworks have made significant progress in natural language processing and computer vision, achieving remarkable performance improvement on various downstream tasks. However, when extended to point cloud data, existing works mainly focus on building ta…

2024

WFTNet: Exploiting Global and Local Periodicity in Long-Term Time Series Forecasting

ICASSP 2024accepted

Recent CNN and Transformer-based models tried to utilize frequency and periodicity information for long-term time series forecasting. However, most existing work is based on Fourier transform, which cannot capture fine-grained and local frequency structure. In this paper, we propose a Wavelet-Fourie…

Cited by 0SourceScholar
2023

Difficulty-Aware Data Augmentor for Scene Text Recognition

ICASSP 2023accepted

Deep neural network (DNN) based scene text recognition (STR) methods usually require a large amount of annotated data for training, which is time-consuming and cost-expensive in practice. To address this issue, many data augmentation methods have been developed to train recognizers by improving the…

Cited by 0SourceScholar
2023

FSR: A General Frequency-Oriented Framework to Accelerate Image Super-resolution Networks

AAAI 2023technical

Deep neural networks (DNNs) have witnessed remarkable achievement in image super-resolution (SR), and plenty of DNN-based SR models with elaborated network designs have recently been proposed. However, existing methods usually require substantial computations by operating in spatial domain. To addre…

2023

Instance-aware Dynamic Prompt Tuning for Pre-trained Point Cloud Models

ICCV 2023poster

Pre-trained point cloud models have found extensive applications in 3D understanding tasks like object classification and part segmentation. However, the prevailing strategy of full fine-tuning in downstream tasks leads to large per-task storage overhead for model parameters, which limits the effici…

Cited by 48PDFcodeScholar
2023

Learned Distributed Image Compression with Multi-Scale Patch Matching in Feature Domain

AAAI 2023technical

Beyond achieving higher compression efficiency over classical image compression codecs, deep image compression is expected to be improved with additional side information, e.g., another image from a different perspective of the same scene. To better utilize the side information under the distributed…

Cited by 13SourcePDFScholar
2023

Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

AAAI 2023technical

Real-world recognition system often encounters the challenge of unseen labels. To identify such unseen labels, multi-label zero-shot learning (ML-ZSL) focuses on transferring knowledge by a pre-trained textual label embedding (e.g., GloVe). However, such methods only exploit single-modal knowledge f…

2023

SFR: Semantic-Aware Feature Rendering of Point Cloud

ICASSP 2023accepted

Multi-view projection methods have demonstrated their ability to reach state-of-the-art performance in point cloud downstream tasks(e.g., classification and retrieval). These methods first require rendering the point cloud into 2D multi-view images. However, conventional methods only project the geo…

Cited by 0SourceScholar
2023

Semantic Preserving Learning for Task-Oriented Point Cloud Downsampling

ICASSP 2023accepted

Recent years have witnessed a tremendous growth in the scale and resolution of point clouds. To facilitate the applications of point cloud in downsampling tasks (e.g., point cloud classification), several task-oriented downsampling works have been developed by training with the task-specific loss wi…

Cited by 0SourceScholar
2023

Towards Robust Scene Text Image Super-resolution via Explicit Location Enhancement

IJCAI 2023poster

Scene text image super-resolution (STISR), aiming to improve image quality while boosting downstream scene text recognition accuracy, has recently achieved great success. However, most existing methods treat the foreground (character regions) and background (non-character regions) equally in the for…

2023

Unsupervised Surface Anomaly Detection with Diffusion Probabilistic Model

ICCV 2023poster

Unsupervised surface anomaly detection aims at discovering and localizing anomalous patterns using only anomaly-free training samples. Reconstruction-based models are among the most popular and successful methods, which rely on the assumption that anomaly regions are more difficult to reconstruct. H…

Cited by 77PDFScholar
2022

Contrastive Quantization with Code Memory for Unsupervised Image Retrieval

AAAI 2022technical

The high efficiency in computation and storage makes hashing (including binary hashing and quantization) a common strategy in large-scale retrieval systems. To alleviate the reliance on expensive annotations, unsupervised deep hashing becomes an important research problem. This paper provides a nove…

2022

NeXT: Towards High Quality Neural Radiance Fields via Multi-Skip Transformer

ECCV 2022poster

"Neural Radiance Fields (NeRF) methods show impressive performance for novel view synthesis by representing a scene via a neural network. However, most existing NeRF based methods, including its variants, treat each sample point individually as input, while ignoring the inherent relationships betwee…

2021

Efficient Face Manipulation Via Deep Feature Disentanglement And Reintegration Net

ICASSP 2021accepted

Deep neural networks (DNNs) have been widely used in facial manipulation. Existing methods focus on training deeper networks in indirect supervision ways (e.g., feature constraint), or in unsupervised ways (e.g., cycle-consistency loss) due to the lack of ground-truth face images for manipulated out…

Cited by 1SourceScholar
2021

HOCA: Higher-Order Channel Attention for Single Image Super-Resolution

ICASSP 2021accepted

Convolutional neural networks (CNNs) have obtained great success in single image super-resolution (SR). More recent works (e.g., RCAN and SAN) have obtained remarkable performance with channel attention based on first- or second-order statistics of features. However, these methods neglect the rich f…

Cited by 5SourceScholar
2021

Knowledge Refinery: Learning from Decoupled Label

AAAI 2021technical

Recently, a variety of regularization techniques have been widely applied in deep neural networks, which mainly focus on the regularization of weight parameters to encourage generalization effectively. Label regularization techniques are also proposed with the motivation of softening the labels whil…

Cited by 15SourcePDFScholar
2019

Hilbert-Based Generative Defense for Adversarial Examples

ICCV 2019poster

Adversarial perturbations of clean images are usually imperceptible for human eyes, but can confidently fool deep neural networks (DNNs) to make incorrect predictions. Such vulnerability of DNNs raises serious security concerns about their practicability in security-sensitive applications. To defend…

Cited by 62PDFScholar
2019

Second-Order Attention Network for Single Image Super-Resolution

CVPR 2019oral

Recently, deep convolutional neural networks (CNNs) have been widely explored in single image super-resolution (SISR) and obtained remarkable performance. However, most of the existing CNN-based SISR methods mainly focus on wider or deeper architecture design, neglecting to explore the feature corre…

Cited by 2042PDFScholar