← Search

Mi Zhang

40 accepted papers

2026

3D-ANC: Adaptive Neural Collapse for Robust 3D Point Cloud Recognition

AAAI 2026technical

Deep neural networks have recently achieved notable progress in 3D point cloud recognition, yet their vulnerability to adversarial perturbations poses critical security challenges in practical deployments. Conventional defense mechanisms struggle to address the evolving landscape of multifaceted att

Cited by 0SourcePDFScholar
2026

DDIM Inversion as a Perturbation Amplifier: Breaking Mimicry Protection via Reconstruction Error Minimization

ICML 2026poster

Personalization techniques for image generation models have increasingly been misused for malicious purposes, including unauthorized style imitation and copyrighted content replication. In response, recent mimicry protection methods embed carefully designed perturbations into images to disrupt a mod…

Cited by 0SourceScholar
2026

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models

CVPR 2026

Vision-language-action (VLA) models unify perception, language, and control for embodied agents but face significant challenges in practical deployment due to rapidly increasing compute and memory demands, especially as models scale to longer horizons and larger backbones. To address these bottlenec

Cited by 0SourcecodeScholar
2026

SafeRoPE: Risk-specific Head-wise Embedding Rotation for Safe Generation in Rectified Flow Transformers

CVPR 2026

Recent Text-to-Image (T2I) models based on rectified-flow transformers (e.g., SD3, FLUX) achieve high generative fidelity but remain vulnerable to unsafe semantics, especially when triggered by multi-token interactions. Existing mitigation methods largely rely on fine-tuning or attention modulation

Cited by 0SourcecodeScholar
2026

SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse

AAAI 2026technical

Despite Video Large Language Models (Video-LLMs) having rapidly advanced in recent years, perceptual hallucinations pose a substantial safety risk, which severely restricts their real-world applicability. While several methods for hallucination mitigation have been proposed, they often compromise th

Cited by 0SourcePDFScholar
2026

Unified Safe In-context Image Generation in Multimodal Diffusion Transformers

ICML 2026poster

Diffusion transformers (DiTs) equipped with multimodal attention (MM-Attn) have become a dominant paradigm for image generation. However, preventing the generation of harmful content remains a critical challenge, particularly in imageto-image (I2I) editing tasks. Existing safety mechanisms are prima…

Cited by 0SourceScholar
2025

$\text{D}_{2}\text{O}$: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models

ICLR 2025poster

Efficient generative inference in Large Language Models (LLMs) is impeded by the growing memory demands of Key-Value (KV) cache, especially for longer sequences. Traditional KV Cache eviction strategies, which discard less critical KV-pairs based on attention scores, often degrade generation quality…

Cited by 0SourcePDFScholar
2025

Argus: Benchmarking and Enhancing Vision-Language Models for 3D Radiology Report Generation

ACL 2025finding

Automatic radiology report generation holds significant potential to streamline the labor-intensive process of report writing by radiologists, particularly for 3D radiographs such as CT scans. While CT scans are critical for clinical diagnostics, they remain less explored compared to 2D radiographs.…

Cited by 0SourcePDFScholar
2025

Creating a Lens of Chinese Culture: A Multimodal Dataset for Chinese Pun Rebus Art Understanding

ACL 2025finding

Large vision-language models (VLMs) have demonstrated remarkable abilities in understanding everyday content. However, their performance in the domain of art, particularly culturally rich art forms, remains less explored. As a pearl of human wisdom and creativity, art encapsulates complex cultural n…

2025

Detect-and-Guide: Self-regulation of Diffusion Models for Safe Text-to-Image Generation via Guideline Token Optimization

CVPR 2025poster

Text-to-image diffusion models have achieved state-of-the-art results in synthesis tasks; however, there is a growing concern about their potential misuse in creating harmful content. To mitigate these risks, post-hoc model intervention techniques, such as concept unlearning and safety guidance, hav…

Cited by 2SourcePDFScholar
2025

InfoCons: Identifying Interpretable Critical Concepts in Point Clouds via Information Theory

ICML 2025poster

Interpretability of point cloud (PC) models becomes imperative given their deployment in safety-critical scenarios such as autonomous vehicles. We focus on attributing PC model outputs to interpretable critical concepts, defined as meaningful subsets of the input point cloud. To enable human-unders…

2025

MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference

NAACL 2025long

Long-context Multimodal Large Language Models (MLLMs) that incorporate long text-image and text-video modalities, demand substantial computational resources as their multimodal Key-Value (KV) cache grows with increasing input lengths, challenging memory and time efficiency. For multimodal scenarios,…

2025

MEIT: Multimodal Electrocardiogram Instruction Tuning on Large Language Models for Report Generation

ACL 2025finding

Electrocardiogram (ECG) is the primary non-invasive diagnostic tool for monitoring cardiac conditions and is crucial in assisting clinicians. Recent studies have concentrated on classifying cardiac conditions using ECG data but have overlooked ECG report generation, which is time-consuming and requi…

2025

Reading Recognition in the Wild

NeurIPS 2025poster

To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of reading recognition to determine when the user is reading. We first introduce the fi…

Cited by 0SourceScholar
2025

SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning

NeurIPS 2025poster

Multimodal large language models (MLLMs) have shown promising capabilities in reasoning tasks, yet still struggle significantly with complex problems requiring explicit self-reflection and self-correction, especially compared to their unimodal text-based counterparts. Existing reflection methods are…

Cited by 0SourceScholar
2025

SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model Compression

NAACL 2025long

Despite significant advancements, the practical deployment of Large Language Models (LLMs) is often hampered by their immense sizes, highlighting the need for effective compression techniques. Singular Value Decomposition (SVD) emerges as a promising method for compressing LLMs. However, existing SV…

2025

SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression

ICLR 2025poster

The advancements in Large Language Models (LLMs) have been hindered by their substantial sizes, which necessitates LLM compression methods for practical deployment. Singular Value Decomposition (SVD) offers a promising solution for LLM compression. However, state-of-the-art SVD-based LLM compression…

2025

The Future Unmarked: Watermark Removal in AI-Generated Images via Next-Frame Prediction

NeurIPS 2025poster

Image watermarking embeds imperceptible signals into AI-generated images for deepfake detection and provenance verification. Although recent semantic-level watermarking methods demonstrate strong resistance against conventional pixel-level removal attacks, their robustness against more advanced remo…

Cited by 0SourceScholar
2024

CausalPC: Improving the Robustness of Point Cloud Classification by Causal Effect Identification

CVPR 2024poster

Deep neural networks have demonstrated remarkable performance in point cloud classification. However previous works show they are vulnerable to adversarial perturbations that can manipulate their predictions. Given the distinctive modality of point clouds various attack strategies have emerged posin…

Cited by 3SourcePDFScholar
2024

ETP: Learning Transferable ECG Representations via ECG-Text Pre-Training

ICASSP 2024accepted

In the domain of cardiovascular healthcare, the Electrocardiogram (ECG) serves as a critical, non-invasive diagnostic tool. Although recent strides in self-supervised learning (SSL) have been promising for ECG representation learning, these techniques often require annotated samples and struggle wit…

Cited by 0SourceScholar
2024

Navigate Beyond Shortcuts: Debiased Learning Through the Lens of Neural Collapse

CVPR 2024highlight

Recent studies have noted an intriguing phenomenon termed Neural Collapse that is when the neural networks establish the right correlation between feature spaces and the training targets their last-layer features together with the classifier weights will collapse into a stable and symmetric structur…

Cited by 6SourcePDFScholar
2023

Black-Box Adversarial Attack on Time Series Classification

AAAI 2023technical

With the increasing use of deep neural network (DNN) in time series classification (TSC), recent work reveals the threat of adversarial attack, where the adversary can construct adversarial examples to cause model mistakes. However, existing researches on the adversarial attack of TSC typically adop…

Cited by 17SourcePDFScholar
2023

CAP: Robust Point Cloud Classification via Semantic and Structural Modeling

CVPR 2023poster

Recently, deep neural networks have shown great success on 3D point cloud classification tasks, which simultaneously raises the concern of adversarial attacks that cause severe damage to real-world applications. Moreover, defending against adversarial examples in point cloud data is extremely diffic…

Cited by 1SourcePDFScholar
2023

FedAudio: A Federated Learning Benchmark for Audio Tasks

ICASSP 2023accepted

Federated learning (FL) has gained substantial attention in recent years due to data privacy concerns related to the pervasiveness of consumer devices that continuously collect data from users. While a number of FL benchmarks have been developed to facilitate FL research, none of them include audio…

Cited by 33SourceScholar
2023

Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing Bias

NeurIPS 2023poster

The scarcity of data presents a critical obstacle to the efficacy of medical vision-language pre-training (VLP). A potential solution lies in the combination of datasets from various language communities. Nevertheless, the main challenge stems from the complexity of integrating diverse syntax and se…

2023

RØROS: Building a Responsive Online Recommender System via Meta-Gradients Updating

ICASSP 2023accepted

In the era of information explosion, users of online services are urgently waiting for timely and effective recommendations. In this paper, we present the first study on the responsiveness aspect of recommender system and present Responsive Online RecOmmender System (RØROS) based on Meta-Gradients U…

Cited by 0SourceScholar
2022

FedRolex: Model-Heterogeneous Federated Learning with Rolling Sub-Model Extraction

NeurIPS 2022accept

Most cross-device federated learning (FL) studies focus on the model-homogeneous setting where the global server model and local client models are identical. However, such constraint not only excludes low-end clients who would otherwise make unique contributions to model training but also restrains…

2022

House of Cans: Covert Transmission of Internal Datasets via Capacity-Aware Neuron Steganography

NeurIPS 2022accept

In this paper, we present a capacity-aware neuron steganography scheme (i.e., Cans) to covertly transmit multiple private machine learning (ML) datasets via a scheduled-to-publish deep neural network (DNN) as the carrier model. Unlike existing steganography schemes which treat the DNN parameters as…

Cited by 2SourcePDFScholar
2022

Multiview Transformers for Video Recognition

CVPR 2022poster

Video understanding requires reasoning at multiple spatiotemporal resolutions -- from short fine-grained motions to events taking place over longer durations. Although transformer architectures have recently advanced the state-of-the-art, they have not explicitly modelled different spatiotemporal re…

Cited by 348PDFcodeScholar
2021

CATE: Computation-aware Neural Architecture Encoding with Transformers

ICML 2021oral

Recent works (White et al., 2020a; Yan et al., 2020) demonstrate the importance of architecture encodings in Neural Architecture Search (NAS). These encodings encode either structure or computation information of the neural architectures. Compared to structure-aware encodings, computation-aware enco…

2021

Dance Revolution: Long-Term Dance Generation with Music via Curriculum Learning

ICLR 2021poster

Dancing to music is one of human's innate abilities since ancient times. In machine learning research, however, synthesizing dance movements from music is a challenging problem. Recently, researchers synthesize human motion sequences through autoregressive models like recurrent neural network (RNN).…

Cited by 159SourcePDFScholar
2020

Does Unsupervised Architecture Representation Learning Help Neural Architecture Search?

NeurIPS 2020poster

Existing Neural Architecture Search (NAS) methods either encode neural architectures using discrete encodings that do not scale well, or adopt supervised learning-based methods to jointly learn architecture representations and optimize architecture search on such representations which incurs search…

2020

MutualNet: Adaptive ConvNet via Mutual Learning from Network Width and Resolution

ECCV 2020poster

We propose the width-resolution mutual learning method (MutualNet) to train a network that is executable at dynamic resource constraints to achieve adaptive accuracy-efficiency trade-offs at runtime. Our method trains a cohort of sub-networks with different widths using different input resolutions t…

2015

Line-Based Multi-Label Energy Optimization for Fisheye Image Rectification and Calibration

CVPR 2015poster

Fisheye image rectification and estimation of intrinsic parameters for real scenes have been addressed in the literature by using line information on the distorted images. In this paper, we propose an easily implemented fisheye image rectification algorithm with line constrains in the undistorted pe…

Cited by 65SourcePDFScholar