← Search

Qing Guo

63 accepted papers

2026

Beyond Visual Reconstruction Quality: Object Perception-aware 3D Gaussian Splatting for Autonomous Driving

ICLR 2026poster

Reconstruction techniques, such as 3D Gaussian Splatting (3DGS), are increasingly used for generating scenarios in autonomous driving system (ADS) research. Existing 3DGS-based works for autonomous driving scenario generation have, through various optimizations, achieved high visual similarity in re…

Cited by 0SourcecodeScholar
2026

DarkDriving: A Real-World Day and Night Aligned Dataset for Autonomous Driving in the Dark Environment

ICRA 2026poster

The low-light conditions are challenging to the vision-centric perception systems for autonomous driving in the dark environment. In this paper, we propose a new benchmark dataset (named DarkDriving) to investigate the low-light enhancement for autonomous driving. The existing real-world low-light e…

2026

Density-Guided Continuous Flow for Robust Counterfactual Explanations

ICML 2026poster

Counterfactual explanations (CEs) are essential for actionable recourse, yet their reliability is often compromised in low-density regions, where classifiers exhibit high variance. Unlike existing methods that rely on expensive ensemble intersections to define stability, we propose DensityFlow, a ge…

Cited by 0SourceScholar
2026

Exploiting Geometric Structures for Modeling Multi-Agent Behaviors: A New Thinking

AAAI 2026technical

In this paper, we rethink model agent behaviors from a geometric structure perspective in multi-agent reinforcement learning. Modeling agent behaviors is essential for understanding how agents interact and facilitating effective decisions. The key lies in capturing the dependencies and sequential re

Cited by 0SourcePDFScholar
2026

FreeMem: Enhancing Consistency in Long Video Generation via Tuning-Free Memory

AAAI 2026technical

Text-to-Video (T2V) generation has advanced greatly, yet maintaining consistency remains challenging, especially for tuning-free long video generation. We attribute the consistency problem to cumulative deviations for long video generation at three levels: the random noise lacking correlation resu

Cited by 0SourcePDFScholar
2026

HOPS: Hierarchical Open-vocabulary Part Segmentation with Attention-Aware Filtering and Affinity-Guided Enhancement

CVPR 2026

Open-vocabulary part segmentation (OVPS) aims to segment objects into fine-grained parts while generalizing to unseen categories. Existing VLM-based methods face two challenges: (1) object over-segmentation, caused by overly broad semantic activations, and (2) part under-segmentation, resulting from

Cited by 0SourcecodeScholar
2026

MAGIC: Mastering Physical Adversarial Generation in Context Through Collaborative LLM Agents

AAAI 2026technical

Physical adversarial attacks in driving scenarios can expose critical vulnerabilities in visual perception models. However, developing such attacks remains non-trivial due to diverse real-world environmental influences. Existing approaches either struggle to generalize to dynamic environments or fai

Cited by 0SourcePDFScholar
2026

Multi-modal Frequency Decomposition Network for Semantic Scene Completion

CVPR 2026

Based on an RGB-D image pair, semantic scene completion (SSC) provides a description for 3D scene understanding by predicting 3D semantic occupancy map. Recent methods extract RGB-D multi-modal features and fuse them in spatial domain, which disregards the misalignment caused by the imperfect raw mu

Cited by 0SourceScholar
2026

OBJVanish: Prompt-Driven Generation of Physically Realizable 3D LiDAR-Invisible Objects

ICML 2026poster

LiDAR-based 3D object detectors are fundamental to autonomous driving, where missed detections pose severe safety risks. While adversarial attacks are crucial for evaluating the robustness of these detectors, existing point-level perturbation methods rarely cause complete object disappearance and pr…

Cited by 0SourceScholar
2026

Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles

CVPR 2026

Most Human-Machine Interaction (HMI) research overlooks the maneuvering needs of passengers in autonomous driving (AD). Natural language offers an intuitive interface, yet translating passenger open-ended instructions into control signals--without sacrificing interpretability and traceability--remai

Cited by 0SourceScholar
2026

PhysPatch: A Physically Realizable and Transferable Adversarial Patch Attack for Multimodal Large Language Models-based Autonomous Driving Systems

AAAI 2026technical

Multimodal Large Language Models (MLLMs) are becoming integral to autonomous driving (AD) systems due to their strong vision-language reasoning capabilities. However, MLLMs are vulnerable to adversarial attacks—particularly adversarial patch attacks—which can pose serious threats in real-world scen

Cited by 0SourcePDFScholar
2026

SURE: Semi-Dense Uncertainty-REfined Feature Matching

ICRA 2026poster

Establishing reliable image correspondences is essential for many robotic vision problems. However, existing methods often struggle in challenging scenarios with large viewpoint changes or textureless regions, where incorrect correspondences may still receive high similarity scores. This is mainly b…

2025

Adaptive Dynamic Programming-Based Fixed-Time Optimal Control for Wheeled Mobile Robot

RA-L 2025

In this study, the adaptive dynamic programming (ADP)-based fixed-time optimal trajectory tracking control is investigated for wheeled mobile robots. An ADP-based fixed-time optimal tracking controller is developed based on the critic-only neural network ADP technique, which guarantees the robot tra

Cited by 5SourceScholar
2025

AngleRoCL: Angle-Robust Concept Learning for Physically View-Invariant Adversarial Patches

NeurIPS 2025poster

Cutting-edge works have demonstrated that text-to-image (T2I) diffusion models can generate adversarial patches that mislead state-of-the-art object detectors in the physical world, revealing detectors' vulnerabilities and risks. However, these methods neglect the T2I patches' attack effectiveness w…

Cited by 0SourcecodeScholar
2025

Concept Matching with Agent for Out-of-Distribution Detection

AAAI 2025technical

The remarkable achievements of Large Language Models (LLMs) have captivated the attention of both academia and industry, transcending their initial role in dialogue generation. To expand the usage scenarios of LLM, some works enhance the effectiveness and capabilities of the model by introducing mor…

2025

Cowpox: Towards the Immunity of VLM-based Multi-Agent Systems

ICML 2025poster

Vision Language Model (VLM) Agents are stateful, autonomous entities capable of perceiving and interacting with their environments through vision and language. Multi-agent systems comprise specialized agents who collaborate to solve a (complex) task. A core security property is **robustness**, stat…

Cited by 0SourcePDFScholar
2025

Defending LVLMs Against Vision Attacks Through Partial-Perception Supervision

ICML 2025poster

Recent studies have raised significant concerns regarding the vulnerability of Large Vision Language Models (LVLMs) to maliciously injected or perturbed input images, which can mislead their responses. Existing defense methods show that such vision attacks are sensitive to image modifications especi…

Cited by 0SourcePDFScholar
2025

DepthVanish: Optimizing Adversarial Interval Structures for Stereo-Depth-Invisible Patches

NeurIPS 2025poster

Stereo depth estimation is a critical task in autonomous driving and robotics, where inaccuracies (such as misidentifying nearby objects as distant) can lead to dangerous situations. Adversarial attacks against stereo depth estimation can help revealing vulnerabilities before deployment. Previous wo…

Cited by 0SourcecodeScholar
2025

Diversifying Robot Locomotion Behaviors with Extrinsic Behavioral Curiosity

ICML 2025poster

Imitation learning (IL) has shown promise in robot locomotion but is often limited to learning a single expert policy, constraining behavior diversity and robustness in unpredictable real-world scenarios. To address this, we introduce Quality Diversity Inverse Reinforcement Learning (QD-IRL), a nove…

2025

Efficient Universal Goal Hijacking with Semantics-guided Prompt Organization

ACL 2025long

Universal goal hijacking is a kind of prompt injection attack that forces LLMs to return a target malicious response for arbitrary normal user prompts. The previous methods achieve high attack performance while being too cumbersome and time-consuming. Also, they have concentrated solely on optimizat…

2025

NoiseController: Towards Consistent Multi-view Video Generation via Noise Decomposition and Collaboration

ICCV 2025poster

High-quality video generation is crucial for many fields, including the film industry and autonomous driving. However, generating videos with spatiotemporal consistencies remains challenging. Current methods typically utilize attention mechanisms or modify noise to achieve consistent videos, neglect…

2025

Open-Vocabulary Part Segmentation via Progressive and Boundary-Aware Strategy

NeurIPS 2025poster

Open-vocabulary part segmentation (OVPS) struggles with structurally connected boundaries due to the inherent conflict between continuous image features and discrete classification mechanism. To address this, we propose PBAPS, a novel training-free framework specifically designed for OVPS. PBAPS lev…

Cited by 0SourcecodeScholar
2025

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

CVPR 2025poster

Large vision-language models (LVLMs) have shown remarkable capabilities in interpreting visual content. While existing works demonstrate these models' vulnerability to deliberately placed adversarial texts, such texts are often easily identifiable as anomalous. In this paper, we present the first ap…

2025

Trajectory-LLM: A Language-based Data Generator for Trajectory Prediction in Autonomous Driving

ICLR 2025poster

Vehicle trajectory prediction is a crucial aspect of autonomous driving, which requires extensive trajectory data to train prediction models to understand the complex, varied, and unpredictable patterns of vehicular interactions. However, acquiring real-world data is expensive, so we advocate using…

2025

VideoShield: Regulating Diffusion-based Video Generation Models via Watermarking

ICLR 2025poster

Artificial Intelligence Generated Content (AIGC) has advanced significantly, particularly with the development of video generation models such as text-to-video (T2V) models and image-to-video (I2V) models. However, like other AIGC types, video generation requires robust content control. A common app…

2024

AdvGPS: Adversarial GPS for Multi-Agent Perception Attack

ICRA 2024poster

The multi-agent perception system collects visual data from sensors located on various agents and leverages their relative poses determined by GPS signals to effectively fuse information, mitigating the limitations of single-agent sensing, such as occlusion. However, the precision of GPS signals can…

Cited by 6SourcecodeScholar
2024

Architecture-Agnostic Iterative Black-Box Certified Defense Against Adversarial Patches

ICASSP 2024accepted

The adversarial patch attack aims to fool image classifiers within a bounded, contiguous region of arbitrary changes. To address this problem in a trustworthy way, the certified patch defense methods are proposed. However, the state-of-the-art certified defenses inevitably needed to access the size…

Cited by 0SourceScholar
2024

ColJailBreak: Collaborative Generation and Editing for Jailbreaking Text-to-Image Deep Generation

NeurIPS 2024poster

The commercial text-to-image deep generation models (e.g. DALL·E) can produce high-quality images based on input language descriptions. These models incorporate a black-box safety filter to prevent the generation of unsafe or unethical content, such as violent, criminal, or hateful imagery. Recent j…

Cited by 2SourcePDFScholar
2024

CosalPure: Learning Concept from Group Images for Robust Co-Saliency Detection

CVPR 2024poster

Co-salient object detection (CoSOD) aims to identify the common and salient (usually in the foreground) regions across a given group of images. Although achieving significant progress state-of-the-art CoSODs could be easily affected by some adversarial perturbations leading to substantial accuracy r…

Cited by 1SourcePDFScholar
2024

Geometry Awakening: Cross-Geometry Learning Exhibits Superiority over Individual Structures

NeurIPS 2024poster

Recent research has underscored the efficacy of Graph Neural Networks (GNNs) in modeling diverse geometric structures within graph data. However, real-world graphs typically exhibit geometrically heterogeneous characteristics, rendering the confinement to a single geometric paradigm insufficient for…

Cited by 0SourcePDFScholar
2024

IRAD: Implicit Representation-driven Image Resampling against Adversarial Attacks

ICLR 2024poster

We introduce a novel approach to counter adversarial attacks, namely, image resampling. Image resampling transforms a discrete image into a new one, simulating the process of scene recapturing or rerendering as specified by a geometrical transformation. The underlying rationale behind our idea is th…

2024

LRR: Language-Driven Resamplable Continuous Representation against Adversarial Tracking Attacks

ICLR 2024poster

Visual object tracking plays a critical role in visual-based autonomous systems, as it aims to estimate the position and size of the object of interest within a live video. Despite significant progress made in this field, state-of-the-art (SOTA) trackers often fail when faced with adversarial pertur…

2024

Light the Night: A Multi-Condition Diffusion Framework for Unpaired Low-Light Enhancement in Autonomous Driving

CVPR 2024poster

Vision-centric perception systems for autonomous driving have gained considerable attention recently due to their cost-effectiveness and scalability especially compared to LiDAR-based systems. However these systems often struggle in low-light conditions potentially compromising their performance and…

Cited by 24SourcePDFScholar
2024

Personalization as a Shortcut for Few-Shot Backdoor Attack against Text-to-Image Diffusion Models

AAAI 2024technical

Although recent personalization methods have democratized high-resolution image synthesis by enabling swift concept acquisition with minimal examples and lightweight computation, they also present an exploitable avenue for highly accessible backdoor attacks. This paper investigates a critical and un…

Cited by 29SourcePDFScholar
2024

SPY-Watermark: Robust Invisible Watermarking for Backdoor Attack

ICASSP 2024accepted

Backdoor attack aims to deceive a victim model when facing backdoor instances while maintaining its performance on benign data. Current methods use manual patterns or special perturbations as triggers, while they often overlook the robustness against data corruption, making backdoor attacks easy to…

Cited by 0SourceScholar
2024

Sim2Real-Fire: A Multi-modal Simulation Dataset for Forecast and Backtracking of Real-world Forest Fire

NeurIPS 2024poster

The latest research on wildfire forecast and backtracking has adopted AI models, which require a large amount of data from wildfire scenarios to capture fire spread patterns. This paper explores using cost-effective simulated wildfire scenarios to train AI models and apply them to the analysis of re…

Cited by 1SourcePDFScholar
2024

Spatial-Temporal Augmentation for Crime Prediction (Student Abstract)

AAAI 2024technical

Crime prediction stands as a pivotal concern within the realm of urban management due to its potential threats to public safety. While prior research has predominantly focused on unraveling the intricate dependencies among urban regions and temporal dynamics, the challenges posed by the scarcity and…

Cited by 1SourcePDFScholar
2024

Voxel Proposal Network via Multi-Frame Knowledge Distillation for Semantic Scene Completion

NeurIPS 2024poster

Semantic scene completion is a difficult task that involves completing the geometry and semantics of a scene from point clouds in a large-scale environment. Many current methods use 3D/2D convolutions or attention mechanisms, but these have limitations in directly constructing geometry and accuratel…

Cited by 1SourcePDFScholar
2023

Background-Mixed Augmentation for Weakly Supervised Change Detection

AAAI 2023technical

Change detection (CD) is to decouple object changes (i.e., object missing or appearing) from background changes (i.e., environment variations) like light and season variations in two images captured in the same scene over a long time span, presenting critical applications in disaster management, urb…

2023

CDUL: CLIP-Driven Unsupervised Learning for Multi-Label Image Classification

ICCV 2023poster

This paper presents a CLIP-based unsupervised learning method for annotation-free multi-label image classification, including three stages: initialization, training, and inference. At the initialization stage, we take full advantage of the powerful CLIP model and propose a novel approach to extend C…

Cited by 47PDFScholar
2023

CMMA: Benchmarking Multi-Affection Detection in Chinese Multi-Modal Conversations

NeurIPS 2023poster

Human communication has a multi-modal and multi-affection nature. The inter-relatedness of different emotions and sentiments poses a challenge to jointly detect multiple human affections with multi-modal clues. Recent advances in this field employed multi-task learning paradigms to render the inter-…

2023

CVSformer: Cross-View Synthesis Transformer for Semantic Scene Completion

ICCV 2023poster

Semantic scene completion (SSC) requires an accurate understanding of the geometric and semantic relationships between the objects in the 3D scene for reasoning the occluded objects. The popular SSC methods voxelize the 3D objects, allowing the deep 3D convolutional network (3D CNN) to learn the obj…

Cited by 9PDFcodeScholar
2023

CopyRNeRF: Protecting the CopyRight of Neural Radiance Fields

ICCV 2023poster

Neural Radiance Fields (NeRF) have the potential to be a major representation of media. Since training a NeRF has never been an easy task, the protection of its model copyright should be a priority. In this paper, by analyzing the pros and cons of possible copyright protection solutions, we propose…

Cited by 43PDFcodeScholar
2023

Distilling Cross-Temporal Contexts for Continuous Sign Language Recognition

CVPR 2023poster

Continuous sign language recognition (CSLR) aims to recognize glosses in a sign language video. State-of-the-art methods typically have two modules, a spatial perception module and a temporal aggregation module, which are jointly learned end-to-end. Existing results in [9,20,25,36] have indicated th…

Cited by 47SourcePDFScholar
2023

Evading DeepFake Detectors via Adversarial Statistical Consistency

CVPR 2023poster

In recent years, as various realistic face forgery techniques known as DeepFake improves by leaps and bounds, more and more DeepFake detection techniques have been proposed. These methods typically rely on detecting statistical differences between natural (i.e., real) and DeepFake-generated images i…

Cited by 59SourcePDFScholar
2023

FAIRER: Fairness as Decision Rationale Alignment

ICML 2023poster

Deep neural networks (DNNs) have made significant progress, but often suffer from fairness issues, as deep models typically show distinct accuracy differences among certain subgroups (e.g., males and females). Existing research addresses this critical issue by employing fairness-aware loss functions…

Cited by 24SourcePDFScholar
2023

Leveraging Inpainting for Single-Image Shadow Removal

ICCV 2023poster

Fully-supervised shadow removal methods achieve the best restoration qualities on public datasets but still generate some shadow remnants. One of the reasons is the lack of large-scale shadow & shadow-free image pairs. Unsupervised methods can alleviate the issue but their restoration qualities are…

Cited by 28PDFcodeScholar
2022

Can You Spot the Chameleon? Adversarially Camouflaging Images From Co-Salient Object Detection

CVPR 2022poster

Co-salient object detection (CoSOD) has recently achieved significant progress and played a key role in retrieval-related tasks. However, it inevitably poses an entirely new safety and security issue, i.e., highly personal and sensitive content can potentially be extracting by powerful CoSOD methods…

Cited by 25PDFcodeScholar
2022

Generative Status Estimation and Information Decoupling for Image Rain Removal

NeurIPS 2022accept

Image rain removal requires the accurate separation between the pixels of the rain streaks and object textures. But the confusing appearances of rains and objects lead to the misunderstanding of pixels, thus remaining the rain streaks or missing the object details in the result. In this paper, we pr…

Cited by 9SourcePDFScholar
2022

MISF: Multi-Level Interactive Siamese Filtering for High-Fidelity Image Inpainting

CVPR 2022poster

Although achieving significant progress, existing deep generative inpainting methods still show low generalization across different scenes. As a result, the generated images usually contain artifacts or the filled pixels differ greatly from the ground truth, making them far from real-world applicati…

Cited by 111PDFcodeScholar
2022

Tight Mutual Information Estimation With Contrastive Fenchel-Legendre Optimization

NeurIPS 2022accept

Successful applications of InfoNCE (Information Noise-Contrastive Estimation) and its variants have popularized the use of contrastive variational mutual information (MI) estimators in machine learning . While featuring superior stability, these estimators crucially depend on costly large-batch trai…

2021

AVA: Adversarial Vignetting Attack against Visual Recognition

IJCAI 2021poster

Vignetting is an inherent imaging phenomenon within almost all optical systems, showing as a radial intensity darkening toward the corners of an image. Since it is a common effect for photography and usually appears as a slight intensity variation, people usually regard it as a part of a photo and w…

Cited by 36SourcePDFScholar
2021

Auto-Exposure Fusion for Single-Image Shadow Removal

CVPR 2021poster

Shadow removal is still a challenging task due to its inherent background-dependent and spatial-variant properties, leading to unknown and diverse shadow patterns. Even powerful deep neural networks could hardly recover traceless shadow-removed background. This paper proposes a new solution for this…

Cited by 174PDFcodeScholar
2021

EfficientDeRain: Learning Pixel-wise Dilation Filtering for High-Efficiency Single-Image Deraining

AAAI 2021technical

Single-image deraining is rather challenging due to the unknown rain model. Existing methods often make specific assumptions of the rain model, which can hardly cover many diverse circumstances in the real world, compelling them to employ complex optimization or progressive refinement. This, however…

2021

Learning To Adversarially Blur Visual Object Tracking

ICCV 2021poster

Motion blur caused by the moving of the object or camera during the exposure can be a key challenge for visual object tracking, affecting tracking accuracy significantly. In this work, we explore the robustness of visual object trackers against motion blur from a new angle, i.e., adversarial blur at…

Cited by 60PDFcodeScholar
2020

SPARK: Spatial-aware Online Incremental Attack Against Visual Tracking

ECCV 2020poster

Adversarial attacks of deep neural networks have been intensively studied on image, audio, natural language, patch, and pixel classification tasks. Nevertheless, as a typical, while important real-world application, the adversarial attacks of online video object tracking that traces an object's movi…

Cited by 112SourcePDFScholar
2020

Watch out! Motion is Blurring the Vision of Your Deep Neural Networks

NeurIPS 2020poster

The state-of-the-art deep neural networks (DNNs) are vulnerable against adversarial examples with additive random-like noise perturbations. While such examples are hardly found in the physical world, the image blurring effect caused by object motion, on the other hand, commonly occurs in practice, m…

2017

Frequency-tuned ACM for biomedical image segmentation

ICASSP 2017accepted

Biomedical images are usually corrupted by strong noise and intensity inhomogeneity simultaneously. Existing region-based active contour models (RACMs) easily fail when segmenting such images. In the frequency domain, we propose a generalized RACM that presents a new way to understand the essence of…

Cited by 0SourceScholar
2017

Learning Dynamic Siamese Network for Visual Object Tracking

ICCV 2017poster

How to effectively learn temporal variation of target appearance, to exclude the interference of cluttered background, while maintaining real-time response, is an essential problem of visual object tracking. Recently, Siamese networks have shown great potentials of matching based trackers in achievi…

Cited by 1046PDFScholar