← Search

Xi Xiao

32 accepted papers

2026

CAD-VAE: Leveraging Correlation-Aware Latents for Comprehensive Fair Disentanglement

AAAI 2026technical

While deep generative models have significantly advanced representation learning, they may inherit or amplify biases and fairness issues by encoding sensitive attributes alongside predictive features. Enforcing strict independence in disentanglement is often unrealistic when target and sensitive fac

Cited by 0SourcePDFScholar
2026

CTR-LORA: CURVATURE-AWARE AND TRUST-REGION GUIDED LOW-RANK ADAPTATION FOR LARGE LANGUAGE MODELS

ICASSP 2026oral

Parameter-efficient fine-tuning (PEFT) has become the standard approach for adapting large language models under limited compute and memory budgets. Although previous methods improve efficiency through low-rank updates, quantization, or heuristic budget reallocation, they often decouple the allocati…

Cited by 0SourcePDFScholar
2026

CrystalDiT: Simple Diffusion Transformers for Crystal Generation

AAAI 2026technical

We present CrystalDiT, a diffusion transformer for crystal structure generation that achieves state-of-the-art performance by challenging the trend of architectural complexity. Instead of intricate, multi-stream designs, CrystalDiT employs a unified transformer that imposes a powerful inductive bias

Cited by 0SourcePDFScholar
2026

Dispersion Loss Counteracts Embedding Condensation and Improves Generalization in Small Language Models

ICML 2026poster

Large language models (LLMs) achieve remarkable performance through ever-increasing parameter counts, but scaling incurs steep computational costs. To better understand LLM scaling, we study representational differences between LLMs and their smaller counterparts, with the goal of replicating the re…

Cited by 0SourceScholar
2026

FeatureFool: Zero-Query Fooling of Video Models via Feature Map

CVPR 2026

The vulnerability of deep neural networks (DNNs) has been preliminarily verified. Existing black-box adversarial attacks usually require multi-round interaction with the model and consume numerous queries, which is impractical in the real-world and hard to scale to recently emerged Video-LLMs. Moreo

Cited by 0SourceScholar
2026

HierAmp: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation

CVPR 2026

Dataset distillation often prioritizes global semantic proximity when creating small surrogate datasets for original large-scale ones. However, object semantics are inherently hierarchical. For example, the position and appearance of a bird's eyes are constrained by the outline of its head. Global p

Cited by 0SourcecodeScholar
2026

Learning Straight Flows: Variational Flow Matching for Efficient Generation

CVPR 2026

Flow Matching has limited ability in achieving one-step generation due to its reliance on learned curved trajectories. Previous studies have attempted to address this limitation by either modifying the coupling distribution to prevent interpolant intersections or introducing consistency and mean-vel

Cited by 0SourceScholar
2026

TWINFUZZ: Dual-Model Fuzzing for Robustness Generalization in Deep Learning

AAAI 2026technical

Deep learning (DL) models are increasingly deployed in safety-critical applications such as face recognition, autonomous driving, and medical diagnosis. Despite their impressive accuracy, they remain vulnerable to adversarial examples - subtle perturbations that can cause incorrect predictions, i.e.

Cited by 0SourcePDFScholar
2026

Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video Approach

AAAI 2026technical

Video-based multimodal large language models (V-MLLMs) have shown vulnerability to adversarial examples in video-text multimodal tasks. However, the transferability of adversarial videos to unseen models—a common and practical real-world scenario—remains unexplored. In this paper, we pioneer an in

Cited by 0SourcePDFScholar
2026

fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding

CVPR 2026

Recent advances in multimodal large language models (LLMs) have enabled unified reasoning across images, audio, and video, but extending such capability to brain imaging remains largely unexplored. Bridging this gap is essential to link neural activity with semantic cognition and to develop cross-mo

Cited by 0SourcecodeScholar
2025

MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization

ICCV 2025poster

Video identity customization seeks to produce high-fidelity videos that maintain consistent identity and exhibit significant dynamics based on users' reference images. However, existing approaches face two key challenges: identity degradation over extended video length and reduced dynamics during tr…

Cited by 0SourcePDFScholar
2025

MoRE-Brain: Routed Mixture of Experts for Interpretable and Generalizable Cross-Subject fMRI Visual Decoding

NeurIPS 2025poster

Decoding visual experiences from fMRI offers a powerful avenue to understand human perception and develop advanced brain-computer interfaces. However, current progress often prioritizes maximizing reconstruction fidelity while overlooking interpretability, an essential aspect for deriving neuroscien…

Cited by 0SourceScholar
2025

Revolutionizing Encrypted Traffic Classification with MH-Net: A Multi-View Heterogeneous Graph Model

AAAI 2025technical

With the growing significance of network security, the classification of encrypted traffic has emerged as an urgent challenge. Traditional byte-based traffic analysis methods are constrained by the rigid granularity of information and fail to fully exploit the diverse correlations between bytes. To…

2025

Sensitivity-LoRA : Low-Load Sensitivity-Based Fine-Tuning for Large Language Models

EMNLP 2025

Large Language Models (LLMs) have transformed both everyday life and scientific research. However, adapting LLMs from general-purpose models to specialized tasks remains challenging, particularly in resource-constrained environments. Low-Rank Adaptation (LoRA), a prominent method within Parameter-Ef

Cited by 0SourcePDFScholar
2025

SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design

EMNLP 2025

Manual slide creation is labor-intensive and requires expert prior knowledge. Existing natural language-based LLM generation methods struggle to capture the visual and structural nuances of slide designs. To address this, we formalize the Reference Image to Slide Generation task and propose Slide2Co

2025

TD-RD: A Top-Down Benchmark with Real-Time Framework for Road Damage Detection

ICASSP 2025accepted

Object detection has witnessed remarkable advancements over the past decade, largely driven by breakthroughs in deep learning and the proliferation of large-scale datasets. However, the domain of road damage detection remains relatively underexplored, despite its critical significance for applicatio…

Cited by 0SourceScholar
2024

Effects of Exponential Gaussian Distribution on (Double Sampling) Randomized Smoothing

ICML 2024poster

Randomized Smoothing (RS) is currently a scalable certified defense method providing robustness certification against adversarial examples. Although significant progress has been achieved in providing defenses against $\ell_p$ adversaries, the interaction between the smoothing distribution and the r…

2024

LogoStyleFool: Vitiating Video Recognition Systems via Logo Style Transfer

AAAI 2024technical

Video recognition systems are vulnerable to adversarial examples. Recent studies show that style transfer-based and patch-based unrestricted perturbations can effectively improve attack efficiency. These attacks, however, face two main challenges: 1) Adding large stylized perturbations to all pixels…

2024

SDformer: Similarity-driven Discrete Transformer For Time Series Generation

NeurIPS 2024poster

The superior generation capabilities of Denoised Diffusion Probabilistic Models (DDPMs) have been effectively showcased across a multitude of domains. Recently, the application of DDPMs has extended to time series generation tasks, where they have significantly outperformed other deep generative mod…

Cited by 6SourcePDFScholar
2024

STEM: Unleashing the Power of Embeddings for Multi-Task Recommendation

AAAI 2024technical

Multi-task learning (MTL) has gained significant popularity in recommender systems as it enables simultaneous optimization of multiple objectives. A key challenge in MTL is negative transfer, but existing studies explored negative transfer on all samples, overlooking the inherent complexities within…

2023

Metis: Understanding and Enhancing In-Network Regular Expressions

NeurIPS 2023poster

Regular expressions (REs) offer one-shot solutions for many networking tasks, e.g., network intrusion detection. However, REs purely rely on expert knowledge and cannot utilize labeled data for better accuracy. Today, neural networks (NNs) have shown superior accuracy and flexibility, thanks to thei…

2022

Efficient and Stable Information Directed Exploration for Continuous Reinforcement Learning

ICASSP 2022accepted

In this paper, we investigate the exploration-exploitation dilemma of reinforcement learning algorithms. We adapt the information directed sampling, an exploration framework that measures the information gain of a policy, to the continuous reinforcement learning. To stabilize the off-policy learning…

Cited by 0SourceScholar
2022

Energy Alignment for Bias Rectification in Class Incremental Learning

ICASSP 2022accepted

In class incremental learning (CIL), models are expected to be able to learn new categories continuously. However, the standard DNNs suffer from catastrophic forgetting. Recent studies show class imbalance is an essential factor that causes catastrophic forgetting in CIL. In this paper, from the per…

Cited by 0SourceScholar
2022

Fine-Tuning Graph Neural Networks via Graph Topology Induced Optimal Transport

IJCAI 2022poster

Recently, the pretrain-finetuning paradigm has attracted tons of attention in graph learning community due to its power of alleviating the lack of labels problem in many real-world applications. Current studies use existing techniques, such as weight constraint, representation constraint, which are…

2020

Adversarial Sparse Transformer for Time Series Forecasting

NeurIPS 2020poster

Many approaches have been proposed for time series forecasting, in light of its significance in wide applications including business demand prediction. However, the existing methods suffer from two key limitations. Firstly, most point prediction models only predict an exact value of each time step…

Cited by 302SourcePDFScholar
2020

Maintaining Discrimination and Fairness in Class Incremental Learning

CVPR 2020poster

Deep neural networks (DNNs) have been applied in class incremental learning, which aims to solve common real-world problems of learning new classes continually. One drawback of standard DNNs is that they are prone to catastrophic forgetting. Knowledge distillation (KD) is a commonly used technique t…

Cited by 614PDFScholar
2020

Self-Paced Probabilistic Principal Component Analysis For Data With Outliers

ICASSP 2020accepted

Principal Component Analysis (PCA) is a popular tool for dimension reduction and feature extraction in data analysis. Probabilistic PCA (PPCA) extends the standard PCA by using a probabilistic model. However, both standard PCA and PPCA are not robust, as they are sensitive to outliers. To alleviate…

Cited by 0SourceScholar
2019

Non-local Self-attention Structure for Function Approximation in Deep Reinforcement Learning

ICASSP 2019accepted

Reinforcement learning is a framework to make sequential decisions. The combination with deep neural networks further improves the ability of this framework. Convolutional nerual networks make it possible to make sequential decisions based on raw pixels information directly and make reinforcement le…

Cited by 0SourceScholar