← Search

Wenqiang Zhang

70 accepted papers

2026

CADiff: Context-Aware Diffusion for Controllable Anomaly Generation in Anomaly Detection

AAAI 2026technical

Generating anomalies is a crucial method to enhance detection and classification performance by expanding anomalous data repository. However, existing anomaly generation methods overlook the intrinsic entanglement between diverse anomaly types and product structures, leading to semantic ambiguity. W

Cited by 0SourcePDFScholar
2026

Cognition-Inspired Dual-Stream Semantic Enhancement for Vision-Based Dynamic Emotion Modeling

ICRA 2026poster

The human brain constructs emotional percepts not by processing facial expressions in isolation, but through a dynamic, hierarchical integration of sensory input with semantic and contextual knowledge. However, existing vision-based dynamic emotion modeling approaches often neglect emotion perceptio…

2026

Commonality in Few: Few-Shot Multimodal Anomaly Detection via Hypergraph-Enhanced Memory

AAAI 2026technical

Few-shot multimodal industrial anomaly detection is a critical yet underexplored task, offering the ability to quickly adapt to complex industrial scenarios. In few-shot settings, insufficient training samples often fail to cover the diverse patterns present in test samples. This challenge can be mi

Cited by 0SourcePDFScholar
2026

GO-PRE:Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D Reconstruction

ICML 2026poster

Active 3D reconstruction relies on active view selection to maximize reconstruction fidelity under limited capture budgets. However, most existing methods rely on surrogate signals—such as parameter uncertainty or geometric heuristics—which are often misaligned with the ultimate goal: the fidelity o…

Cited by 0SourceScholar
2026

LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have shown great promise but require substantial computational resources during inference. Attackers can exploit this by inducing excessive output, leading to resource exhaustion and service degradation. Prior energy-latency attacks aim to increase generation…

Cited by 0SourceScholar
2026

PositionIC: Unified Position and Identity Consistency for Image Customization

CVPR 2026

Recent subject-driven image customization excels in fidelity, yet fine-grained instance-level spatial control remains an elusive challenge, hindering real-world applications. This limitation stems from two factors: a scarcity of scalable, position-annotated datasets, and the entanglement of identity

Cited by 0SourcecodeScholar
2026

RSAgent: Learning to Reason and Act via Multi-Turn Tool Invocations for Text-Guided Segmentation

ICML 2026poster

Text-guided object segmentation requires both cross-modal reasoning and pixel grounding abilities. Most recent methods treat it as a single forward pass, where the model directly predicts pixel prompts to a segmentation model, which limits verification, refocusing and refinement when initial localiz…

Cited by 0SourceScholar
2026

SURF-Loco: Mastering Complex Industrial Terrains with 3D Surfel-Based Reinforcement Learning for Legged Robots

ICRA 2026poster

Legged robots offer significant potential for navigating complex industrial terrains, but their capabilities are often constrained by perception systems struggling to interpret intricate 3D geometry. Conventional 2D/2.5D representations like depth or elevation maps fail to capture complex 3D geometr…

Cited by 0Scholar
2026

Seeing Is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding

AAAI 2026technical

Multimodal Large Language Models (MLLMs) have unlocked powerful cross-modal capabilities, but still significantly suffer from hallucinations. As such, accurate detection of hallucinations in MLLMs is imperative for ensuring their reliability in practical applications. To this end, guided by the prin

Cited by 0SourcePDFScholar
2026

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts

ICML 2026poster

Multimodal visual object tracking can be divided into to several kinds of tasks (e.g. RGB and RGB+X tracking), based on the input modality. Existing methods often train separate models for each modality or rely on pretrained models to adapt to new modalities, which limits efficiency, scalability, an…

Cited by 0SourceScholar
2025

Agentic RL Scaling Law: Spontaneous Code Execution for Mathematical Problem Solving

NeurIPS 2025poster

Large Language Models (LLMs) often struggle with mathematical reasoning tasks requiring precise, verifiable computation. While Reinforcement Learning (RL) from outcome-based rewards enhances text-based reasoning, understanding how agents autonomously learn to leverage external tools like code execu…

Cited by 0SourcecodeScholar
2025

Boosting Adversarial Transferability with Spatial Adversarial Alignment

NeurIPS 2025poster

Deep neural networks are vulnerable to adversarial examples that exhibit transferability across various models. Numerous approaches are proposed to enhance the transferability of adversarial examples, including advanced optimization, data augmentation, and model modifications. However, these methods…

Cited by 0SourceScholar
2025

Component-Aware Unsupervised Logical Anomaly Generation for Industrial Anomaly Detection

ICRA 2025

Anomaly detection is critical in industrial manufacturing for ensuring product quality and improving efficiency in automated processes. The scarcity of anomalous samples limits traditional detection methods, making anomaly generation essential for expanding the data repository. However, recent gener

Cited by 2SourceScholar
2025

D2SP: Dynamic Dual-Stage Purification Framework for Dual Noise Mitigation in Vision-based Affective Recognition.

CVPR 2025poster

The current advancements in Dynamic Facial Expression Recognition (DFER) methods mainly focus on better capturing the spatial and temporal features of facial expressions. However, DFER datasets contain a substantial amount of noisy samples, and few have addressed the issue of handling this noise. We…

Cited by 0SourcePDFScholar
2025

Dynamic Semantic-Aware Correlation Modeling for UAV Tracking

NeurIPS 2025poster

UAV tracking can be widely applied in scenarios such as disaster rescue, environmental monitoring, and logistics transportation. However, existing UAV tracking methods predominantly emphasize speed and lack exploration in semantic awareness, which hinders the search region from extracting accurate l…

Cited by 0SourceScholar
2025

Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences Alignment

NeurIPS 2025poster

Preference alignment in diffusion models has primarily focused on benign human preferences (e.g., aesthetic). In this paper, we propose a novel perspective: framing unrestricted adversarial example generation as a problem of aligning with adversary preferences. Unlike benign alignment, adversarial a…

Cited by 0SourceScholar
2025

ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation

NeurIPS 2025poster

Vision-Language-Action (VLA) models have advanced general-purpose robotic manipulation by leveraging pretrained visual and linguistic representations. However, they struggle with contact-rich tasks that require fine-grained control involving force, especially under visual occlusion or dynamic uncert…

Cited by 0SourceScholar
2025

General Compression Framework for Efficient Transformer Object Tracking

ICCV 2025poster

Previous works have attempted to improve tracking efficiency through lightweight architecture design or knowledge distillation from teacher models to compact student trackers. However, these solutions often sacrifice accuracy for speed to a great extent, and also have the problems of complex trainin…

2025

Noise Fusion-based Distillation Learning for Anomaly Detection in Complex Industrial Environments

IROS 2025

Anomaly detection and localization in automated industrial manufacturing can significantly enhance production efficiency and product quality. Existing methods are capable of detecting surface defects in pre-defined or controlled imaging environments. However, accurately detecting workpiece defects i

Cited by 0SourcecodeScholar
2025

OUS: Bridging Scene Context and Facial Features to Overcome the Rigid Cognitive Problem

AAAI 2025technical

Dynamic Facial Expression Recognition (DFER) is crucial for affective computing but often overlooks the impact of scene context. We have identified a significant issue in current DFER tasks: human annotators typically integrate emotions from various angles, including environmental cues and body lang…

2025

OpenVIS: Open-vocabulary Video Instance Segmentation

AAAI 2025technical

Open-vocabulary Video Instance Segmentation (OpenVIS) can simultaneously detect, segment, and track arbitrary object categories in a video, without being constrained to categories seen during training. In this work, we propose InstFormer, a carefully designed framework for the OpenVIS task that achi…

2025

Scoring, Remember, and Reference: Catching Camouflaged Objects in Videos

ICCV 2025poster

Video Camouflaged Object Detection (VCOD) aims to segment objects whose appearances closely resemble their surroundings, posing a challenging and emerging task. Existing vision models often struggle in such scenarios due to the indistinguishable appearance of camouflaged objects and the insufficient…

Cited by 0SourcePDFScholar
2025

Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection

ICCV 2025poster

Pre-trained vision-language models have exhibited remarkable abilities in detecting out-of-distribution (OOD) samples. However, some challenging OOD samples, which lie close to in-distribution (InD) data in image feature space, can still lead to misclassification. The emergence of foundation models…

2024

A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis

AAAI 2024technical

Well-designed prompts have demonstrated the potential to guide text-to-image models in generating amazing images. Although existing prompt engineering methods can provide high-level guidance, it is challenging for novice users to achieve the desired results by manually entering prompts due to a disc…

2024

Adaptive Multi-modal Fusion of Spatially Variant Kernel Refinement with Diffusion Model for Blind Image Super-Resolution

ECCV 2024poster

"Pre-trained diffusion models utilized for image generation encapsulate a substantial reservoir of a priori knowledge pertaining to intricate textures. Harnessing the potential of leveraging this a priori knowledge in the context of image super-resolution presents a compelling avenue. Nonetheless, p…

Cited by 3SourcePDFScholar
2024

De-confounded Data-free Knowledge Distillation for Handling Distribution Shifts

CVPR 2024poster

Data-Free Knowledge Distillation (DFKD) is a promising task to train high-performance small models to enhance actual deployment without relying on the original training data. Existing methods commonly avoid relying on private data by utilizing synthetic or sampled data. However a long-overlooked iss…

Cited by 6SourcePDFScholar
2024

DeTrack: In-model Latent Denoising Learning for Visual Object Tracking

NeurIPS 2024poster

Previous visual object tracking methods employ image-feature regression models or coordinate autoregression models for bounding box prediction. Image-feature regression methods heavily depend on matching results and do not utilize positional prior, while the autoregressive approach can only be train…

Cited by 0SourcePDFScholar
2024

FD-UAD: Unsupervised Anomaly Detection Platform Based on Defect Autonomous Imaging and Enhancement

IJCAI 2024poster

In industrial quality control, detecting defects is essential. However, manual checks and machine vision encounter challenges in complex conditions, as defects vary among products made of different materials and shapes. We create FD-UAD, Unsupervised Anomaly Detection Platform Based on Defect Autono…

Cited by 0SourcePDFScholar
2024

LCGen: Mining in Low-Certainty Generation for View-consistent Text-to-3D

NeurIPS 2024poster

The Janus Problem is a common issue in SDS-based text-to-3D methods. Due to view encoding approach and 2D diffusion prior guidance, the 3D representation model tends to learn content with higher certainty from each perspective, leading to view inconsistency. In this work, we first model and analyze…

Cited by 0SourcePDFScholar
2024

MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution

NeurIPS 2024poster

In software development, resolving the emergent issues within GitHub repositories is a complex challenge that involves not only the incorporation of new code but also the maintenance of existing code. Large Language Models (LLMs) have shown promise in code generation but face difficulties in resolvi…

Cited by 39SourcePDFScholar
2024

OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning

CVPR 2024highlight

Visual object tracking aims to localize the target object of each frame based on its initial appearance in the first frame. Depending on the input modility tracking tasks can be divided into RGB tracking and RGB+X (e.g. RGB+N and RGB+D) tracking. Despite the different input modalities the core aspec…

Cited by 62SourcePDFScholar
2024

Out of Thin Air: Exploring Data-Free Adversarial Robustness Distillation

AAAI 2024technical

Adversarial Robustness Distillation (ARD) is a promising task to solve the issue of limited adversarial robustness of small capacity models while optimizing the expensive computational costs of Adversarial Training (AT). Despite the good robust performance, the existing ARD methods are still impract…

Cited by 9SourcePDFScholar
2024

PanoVOS: Bridging Non-panoramic and Panoramic Views with Transformer for Video Segmentation

ECCV 2024poster

"Panoramic videos contain richer spatial information and have attracted tremendous amounts of attention due to their exceptional experience in some fields such as autonomous driving and virtual reality. However, existing datasets for video segmentation only focus on conventional planar images. To ad…

2024

Pixel-level Semantic Correspondence through Layout-aware Representation Learning and Multi-scale Matching Integration

CVPR 2024poster

Establishing precise semantic correspondence across object instances in different images is a fundamental and challenging task in computer vision. In this task difficulty arises often due to three challenges: confusing regions with similar appearance inconsistent object scale and indistinguishable n…

2023

Adversarial Contrastive Distillation with Adaptive Denoising

ICASSP 2023accepted

Adversarial Robustness Distillation (ARD) is a novel method to boost the robustness of small models. Unlike general adversarial training, its robust knowledge transfer can be less easily restricted by the model capacity. However, the teacher model that provides the robustness of knowledge does not a…

Cited by 0SourceScholar
2023

CiCo: Domain-Aware Sign Language Retrieval via Cross-Lingual Contrastive Learning

CVPR 2023poster

This work focuses on sign language retrieval--a recently proposed task for sign language understanding. Sign language retrieval consists of two sub-tasks: text-to-sign-video (T2V) retrieval and sign-video-to-text (V2T) retrieval. Different from traditional video-text retrieval, sign language videos,…

2023

Content-based Unrestricted Adversarial Attack

NeurIPS 2023poster

Unrestricted adversarial attacks typically manipulate the semantic content of an image (e.g., color or texture) to create adversarial examples that are both effective and photorealistic, demonstrating their ability to deceive human perception and deep neural networks with stealth and success. Howeve…

Cited by 86SourcePDFScholar
2023

Correspondence Transformers With Asymmetric Feature Learning and Matching Flow Super-Resolution

CVPR 2023poster

This paper solves the problem of learning dense visual correspondences between different object instances of the same category with only sparse annotations. We decompose this pixel-level semantic matching problem into two easier ones: (i) First, local feature descriptors of source and target images…

2023

Efficient Decision-based Black-box Patch Attacks on Video Recognition

ICCV 2023poster

Although Deep Neural Networks (DNNs) have demonstrated excellent performance, they are vulnerable to adversarial patches that introduce perceptible and localized perturbations to the input. Generating adversarial patches on images has received much attention, while adversarial patches on videos have…

Cited by 23PDFScholar
2023

Hierarchical Visual Categories Modeling: A Joint Representation Learning and Density Estimation Framework for Out-of-Distribution Detection

ICCV 2023poster

Detecting out-of-distribution inputs for visual recognition models has become critical in safe deep learning. This paper proposes a novel hierarchical visual category modeling scheme to separate out-of-distribution data from in-distribution data through joint representation learning and statistical…

Cited by 3PDFScholar
2023

Improving Generalization in Visual Reinforcement Learning via Conflict-aware Gradient Agreement Augmentation

ICCV 2023poster

Learning a policy with great generalization to unseen environments remains challenging but critical in visual reinforcement learning. Despite the success of augmentation combination in the supervised learning generalization, naively applying it to visual RL algorithms may damage the training efficie…

Cited by 25PDFScholar
2023

LVOS: A Benchmark for Long-term Video Object Segmentation

ICCV 2023poster

Existing video object segmentation (VOS) benchmarks focus on short-term videos which just last about 3-5 seconds and where objects are visible most of the time. These videos are poorly representative of practical applications, and the absence of long-term datasets restricts further investigation of…

Cited by 60PDFcodeScholar
2023

MISC210K: A Large-Scale Dataset for Multi-Instance Semantic Correspondence

CVPR 2023poster

Semantic correspondence have built up a new way for object recognition. However current single-object matching schema can be hard for discovering commonalities for a category and far from the real-world recognition tasks. To fill this gap, we design the multi-instance semantic correspondence task wh…

2023

RankDNN: Learning to Rank for Few-Shot Learning

AAAI 2023technical

This paper introduces a new few-shot learning pipeline that casts relevance ranking for image retrieval as binary ranking relation classification. In comparison to image classification, ranking relation classification is sample efficient and domain agnostic. Besides, it provides a new perspective on…

2023

Reading Relevant Feature from Global Representation Memory for Visual Object Tracking

NeurIPS 2023poster

Reference features from a template or historical frames are crucial for visual object tracking. Prior works utilize all features from a fixed template or memory for visual object tracking. However, due to the dynamic nature of videos, the required reference historical information for different searc…

Cited by 16SourcePDFScholar
2023

TINYCOD: Tiny and Effective Model for Camouflaged Object Detection

ICASSP 2023accepted

This paper introduces an effective and tiny model for real-time Camouflaged Object Detection (COD) named Tiny-COD. It achieves high performance with very low costs (Parameters < 5M, FLOPs < 1.5G), which can be applied on mobile devices. Specifically, we introduce a simple but effective Adjacent Scal…

Cited by 0SourceScholar
2022

Attribute Surrogates Learning and Spectral Tokens Pooling in Transformers for Few-Shot Learning

CVPR 2022poster

This paper presents new hierarchically cascaded transformers that can improve data efficiency through attribute surrogates learning and spectral tokens pooling. Vision transformers have recently been thought of as a promising alternative to convolutional neural networks for visual recognition. But w…

Cited by 71PDFcodeScholar
2022

AziNorm: Exploiting the Radial Symmetry of Point Cloud for Azimuth-Normalized 3D Perception

CVPR 2022poster

Studying the inherent symmetry of data is of great importance in machine learning. Point cloud, the most important data format for 3D environmental perception, is naturally endowed with strong radial symmetry. In this work, we exploit this radial symmetry via a divide-and-conquer strategy to boost 3…

Cited by 7PDFcodeScholar
2022

Efficient Universal Shuffle Attack for Visual Object Tracking

ICASSP 2022accepted

Recently, adversarial attacks have been applied in visual object tracking to deceive deep trackers by injecting imperceptible perturbations into video frames. However, previous work only generates the video-specific perturbations, which restricts its application scenarios. In addition, existing atta…

Cited by 0SourceScholar
2022

FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos

CVPR 2022poster

Current benchmarks for facial expression recognition (FER) mainly focus on static images, while there are limited datasets for FER in videos. It is still ambiguous to evaluate whether performances of existing methods remain satisfactory in real-world application-oriented scenes. For example, the "Ha…

Cited by 110PDFcodeScholar
2022

Position-aware Joint Entity and Relation Extraction with Attention Mechanism

IJCAI 2022poster

Named entity recognition and relation extraction are two important core subtasks of information extraction, which aim to identify named entities and extract relations between them. In recent years, span representation methods have received a lot of attention and are widely used to extract entities a…

Cited by 6SourcePDFScholar
2022

Selective Scale Cascade Attention Network for Breast Cancer Histopathology Image Classification

ICASSP 2022accepted

Convolutional Neural Networks (CNNs) approaches are widely applied to histopathological image analysis due to the breakthrough performance achieved. However, it remains challenging because complex backgrounds obscure the most discriminative region. In this paper, we propose selective scale cascade a…

Cited by 0SourceScholar
2022

Sparse Instance Activation for Real-Time Instance Segmentation

CVPR 2022poster

In this paper, we propose a conceptually novel, efficient, and fully convolutional framework for real-time instance segmentation. Previously, most instance segmentation methods heavily rely on object detection and perform mask prediction based on bounding boxes or dense centers. In contrast, we prop…

Cited by 182PDFcodeScholar
2022

TopFormer: Token Pyramid Transformer for Mobile Semantic Segmentation

CVPR 2022poster

Although vision transformers (ViTs) have achieved great success in computer vision, the heavy computational cost hampers their applications to dense prediction tasks such as semantic segmentation on mobile devices. In this paper, we present a mobile-friendly architecture named Token Pyramid Vision T…

Cited by 304PDFcodeScholar
2022

Towards Practical Certifiable Patch Defense With Vision Transformer

CVPR 2022poster

Patch attacks, one of the most threatening forms of physical attack in adversarial examples, can lead networks to induce misclassification by modifying pixels arbitrarily in a continuous region. Certifiable patch defense can guarantee robustness that the classifier is not affected by patch attacks.…

Cited by 81PDFScholar
2022

Weakly-Supervised Salient Object Detection Using Point Supervision

AAAI 2022technical

Current state-of-the-art saliency detection models rely heavily on large datasets of accurate pixel-wise annotations, but manually labeling pixels is time-consuming and labor-intensive. There are some weakly supervised methods developed for alleviating the problem, such as image label, bounding box…

2021

Dual Path Learning for Domain Adaptation of Semantic Segmentation

ICCV 2021poster

Domain adaptation for semantic segmentation enables to alleviate the need for large-scale pixel-wise annotations. Recently, self-supervised learning (SSL) with a combination of image-to-image translation shows great effectiveness in adaptive segmentation. The most common practice is to perform SSL a…

Cited by 82PDFcodeScholar
2021

Dual-Stream Network Based On Global Guidance for Salient Object Detection

ICASSP 2021accepted

High-level features can help low-level features eliminate semantic ambiguity, which is crucial for obtaining the precise salient object. Some methods use high-level features to provide global guidance for some layers of the network. However, there remain several problems: (1) the global guidance has…

Cited by 0SourceScholar
2021

Improving Zero-Shot Cross-lingual Transfer for Multilingual Question Answering over Knowledge Graph

NAACL 2021long

Multilingual question answering over knowledge graph (KGQA) aims to derive answers from a knowledge graph (KG) for questions in multiple languages. To be widely applicable, we focus on its zero-shot transfer setting. That is, we can only access training data in a high-resource language, while need t…

2020

All In One Network for Driver Attention Monitoring

ICASSP 2020accepted

Nowadays, driver drowsiness and driver distraction is considered as a major risk for fatal road accidents around the world. As a result, driver monitoring identifying is emerging as an essential function of automotive safety systems. Its basic features include head pose, gaze direction, yawning and…

Cited by 0SourceScholar
2020

Multi-Scale Deep Feature Fusion for Vehicle Re-Identification

ICASSP 2020accepted

Vehicle re-identification (re-id) is challenging due to the small inter-class distance. The differences between similar vehicles can be extremely subtle and only captured at particular scales and semantic levels. In this paper, we propose a novel Multi-Scale Deep Feature Fusion Network (MSDeep) to c…

Cited by 0SourceScholar
2018

MetaAnchor: Learning to Detect Objects with Customized Anchors

NeurIPS 2018poster

We propose a novel and flexible anchor mechanism named MetaAnchor for object detection frameworks. Unlike many previous detectors model anchors via a predefined manner, in MetaAnchor anchor functions could be dynamically generated from the arbitrary customized prior boxes. Taking advantage of weight…