← Search

Ziqi Zhou

22 accepted papers

2026

UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models

CVPR 2026

With the advancement of multi-modal Large Language Models (LLMs), Video LLMs have been further developed to perform on holistic and specialized video understanding. However, existing works are limited to specialized video understanding tasks, failing to achieve a comprehensive and multi-grained vide

Cited by 0SourceScholar
2025

AdvEDM: Fine-grained Adversarial Attack against VLM-based Embodied Agents

NeurIPS 2025poster

Vision-Language Models (VLMs), with their strong reasoning and planning capabilities, are widely used in embodied decision-making (EDM) tasks in embodied agents, such as autonomous driving and robotic manipulation. Recent research has increasingly explored adversarial attacks on VLMs to reveal their…

Cited by 0SourceScholar
2025

BadRobot: Jailbreaking Embodied LLM Agents in the Physical World

ICLR 2025poster

Embodied AI represents systems where AI is integrated into physical entities. Multimodal Large Language Model (LLM), which exhibits powerful language understanding abilities, has been extensively employed in embodied AI by facilitating sophisticated task planning. However, a critical safety issue re…

Cited by 0SourcePDFScholar
2025

Breaking Barriers in Physical-World Adversarial Examples: Improving Robustness and Transferability via Robust Feature

AAAI 2025technical

As deep neural networks (DNNs) are widely applied in the physical world, many researches are focusing on physical-world adversarial examples (PAEs), which introduce perturbations to inputs and cause the model's incorrect outputs. However, existing PAEs face two challenges: unsatisfactory attack perf…

2025

Detecting and Corrupting Convolution-based Unlearnable Examples

AAAI 2025technical

Convolution-based unlearnable examples (UEs) employ class-wise multiplicative convolutional noise to training samples, severely compromising model performance. This fire-new type of UEs have successfully countered all defense mechanisms against UEs. The failure of such defenses can be attributed to…

2025

Diffused Poses and Distilled Expressions for Controllable Audio-driven Talking Face Generation

ICASSP 2025accepted

Audio-driven portrait animation is an emerging field in multi-modal generation that aims to create lifelike talking face videos from audio input. While significant progress has been made, accurately modeling the relationship between audio signals and various facial motions, such as head poses and ex…

Cited by 0SourceScholar
2025

GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic Expressions

AAAI 2025technical

Audio-driven talking head generation necessitates seamless integration of audio and visual data amidst the challenges posed by diverse input portraits and intricate correlations between audio and facial motions. In response, we propose a robust framework GoHD designed to produce highly realistic, ex…

2025

MARS: A Malignity-Aware Backdoor Defense in Federated Learning

NeurIPS 2025poster

Federated Learning (FL) is a distributed paradigm aimed at protecting participant data privacy by exchanging model parameters to achieve high-quality model training. However, this distributed nature also makes FL highly vulnerable to backdoor attacks. Notably, the recently proposed state-of-the-art…

Cited by 0SourceScholar
2025

NumbOD: A Spatial-Frequency Fusion Attack Against Object Detectors

AAAI 2025technical

With the advancement of deep learning, object detectors (ODs) with various architectures have achieved significant success in complex scenarios like autonomous driving. Previous adversarial attacks against ODs have been focused on designing customized attacks targeting their specific structures (eg,…

2025

PB-UAP: Hybride Universal Adversarial Attack for Image Segmentation

ICASSP 2025accepted

With the rapid advancement of deep learning, the model robustness has become a significant research hotspot, i.e., adversarial attacks on deep neural networks. Existing works primarily focus on image classification tasks, aiming to alter the model’s predicted labels. Due to the output complexity and…

Cited by 0SourceScholar
2025

Test-Time Backdoor Detection for Object Detection Models

CVPR 2025poster

Object detection models are vulnerable to backdoor attacks, where attackers poison a small subset of training samples by embedding a predefined trigger to manipulate prediction. Detecting poisoned samples (i.e., those containing triggers) at test time can prevent backdoor activation. However, unlike…

Cited by 1SourcePDFScholar
2025

Vanish into Thin Air: Cross-prompt Universal Adversarial Attacks for SAM2

NeurIPS 2025spotlight

Recent studies reveal the vulnerability of the image segmentation foundation model SAM to adversarial examples. Its successor, SAM2, has attracted significant attention due to its strong generalization capability in video segmentation. However, its robustness remains unexplored, and it is unclear wh…

Cited by 0SourceScholar
2025

Visual Self-Refinement for Autoregressive Models

EMNLP 2025

Autoregressive models excel in sequential modeling and have proven to be effective for vision-language data. However, the spatial nature of visual signals conflicts with the sequential dependencies of next-token prediction, leading to suboptimal results. This work proposes a plug-and-play refinement

Cited by 0SourcePDFScholar
2024

A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties

ECCV 2024poster

"We introduce ProLab, a novel approach using property-level label space for creating strong interpretable segmentation models. Instead of relying solely on category-specific annotations, ProLab uses descriptive properties grounded in common sense knowledge for supervising segmentation models. It is…

2024

DarkSAM: Fooling Segment Anything Model to Segment Nothing

NeurIPS 2024poster

Segment Anything Model (SAM) has recently gained much attention for its outstanding generalization to unseen data and tasks. Despite its promising prospect, the vulnerabilities of SAM, especially to universal adversarial perturbation (UAP) have not been thoroughly investigated yet. In this paper, we…

2024

Detector Collapse: Backdooring Object Detection to Catastrophic Overload or Blindness in the Physical World

IJCAI 2024poster

Object detection tasks, crucial in safety-critical systems like autonomous driving, focus on pinpointing object locations. These detectors are known to be susceptible to backdoor attacks. However, existing backdoor techniques have primarily been adapted from classification tasks, overlooking deeper…

Cited by 13SourcePDFScholar
2024

Unlearnable 3D Point Clouds: Class-wise Transformation Is All You Need

NeurIPS 2024poster

Traditional unlearnable strategies have been proposed to prevent unauthorized users from training on the 2D image data. With more 3D point cloud data containing sensitivity information, unauthorized usage of this new type data has also become a serious concern. To address this, we propose the first…

2023

Downstream-agnostic Adversarial Examples

ICCV 2023poster

Self-supervised learning usually uses a large amount of unlabeled data to pre-train an encoder which can be used as a general-purpose feature extractor, such that downstream users only need to perform fine-tuning operations to enjoy the benefit of "big model". Despite this promising prospect, the se…

Cited by 30PDFcodeScholar
2022

Generalizable Cross-Modality Medical Image Segmentation via Style Augmentation and Dual Normalization

CVPR 2022poster

For medical image segmentation, imagine if a model was only trained using MR images in source domain, how about its performance to directly segment CT images in target domain? This setting, namely generalizable cross-modality segmentation, owning its clinical potential, is much more challenging than…

Cited by 90PDFcodeScholar
2022

Generalizable Medical Image Segmentation via Random Amplitude Mixup and Domain-Specific Image Restoration

ECCV 2022poster

"For medical image analysis, segmentation models trained on one or several domains lack generalization ability to unseen domains due to discrepancies between different data acquisition policies. We argue that the degeneration in segmentation performance is mainly attributed to overfitting to source…

2021

Binocular Mutual Learning for Improving Few-Shot Classification

ICCV 2021poster

Most of the few-shot learning methods learn to transfer knowledge from datasets with abundant labeled data (i.e., the base set). From the perspective of class space on base set, existing methods either focus on utilizing all classes under a global view by normal pretraining, or pay more attention to…

Cited by 113PDFcodeScholar