← Search

Hang Wang

26 accepted papers

2026

QRShield: Exploiting Vulnerabilities of Latent Diffusion Models for Preventing AI Art Plagiarism

AAAI 2026technical

Latent Diffusion Models (LDMs) have achieved remarkable success in image generation tasks, yet their low barrier to customization poses severe threats related to art plagiarism. As a countermeasure, adversarial methods have been proposed to protect artworks from plagiarism. However, current methods

Cited by 0SourcePDFScholar
2026

VITA: Vision-to-Action Flow Matching Policy

ICLR 2026poster

Conventional flow matching and diffusion-based policies sample through iterative denoising from standard noise distributions (e.g., Gaussian), and require conditioning modules to repeatedly incorporate visual information during the generative process, incurring substantial time and memory overhead.…

Cited by 0SourcecodeScholar
2025

AdaWM: Adaptive World Model based Planning for Autonomous Driving

ICLR 2025poster

World model based reinforcement learning (RL) has emerged as a promising approach for autonomous driving, which learns a latent dynamics model and uses it to train a planning policy. To speed up the learning process, the pretrain-finetune paradigm is often used, where online RL is initialized by a…

Cited by 1SourcePDFScholar
2025

One-Shot Face Avatar Generation in a Single Forward Pass with Identity Preservation

ICASSP 2025accepted

Face avatar generation has gained significant attention recently. With the help of the Neural Radiance Field (NeRF), existing 3D methods alleviate facial distortion in 2D methods under large pose changes. However, the state-of-the-art 3D methods still require additional optimization for generation o…

Cited by 0SourceScholar
2024

A Multi-Scale Convolutional Hybrid Attention Residual Network for Enhancing Underwater Image and Identifying Underwater Multi-Scene Sea Cucumber

RA-L 2024

At present, the use of underwater robots to replace underwater manual work is a future development direction. The complex and changeable underwater environment brings great difficulties to the operation of robots. In order to improve the problem of color distortion and degradation of sea cucumber im

Cited by 2SourceScholar
2024

Breaking Semantic Artifacts for Generalized AI-generated Image Detection

NeurIPS 2024poster

With the continuous evolution of AI-generated images, the generalized detection of them has become a crucial aspect of AI security. Existing detectors have focused on cross-generator generalization, while it remains unexplored whether these detectors can generalize across different image scenes, e.…

2024

Intrinsic Phase-Preserving Networks for Depth Super Resolution

AAAI 2024technical

Depth map super-resolution (DSR) plays an indispensable role in 3D vision. We discover an non-trivial spectral phenomenon: the components of high-resolution (HR) and low-resolution (LR) depth maps manifest the same intrinsic phase, and the spectral phase of RGB is a superset of them, which suggests…

2024

Object-Oriented Anchoring and Modal Alignment in Multimodal Learning

ECCV 2024poster

"Modality alignment has been of paramount importance in recent developments of multimodal learning, which has inspired many innovations in multimodal networks and pre-training tasks. Single-stream networks can effectively leverage self-attention mechanisms to facilitate modality interactions but suf…

Cited by 1SourcePDFScholar
2024

Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis

IJCAI 2024poster

Detecting synthetic from real speech is increasingly crucial due to the risks of misinformation and identity impersonation. While various datasets for synthetic speech analysis have been developed, they often focus on specific areas, limiting their utility for comprehensive research. To fill this ga…

2023

Deep Arbitrary-Scale Image Super-Resolution via Scale-Equivariance Pursuit

CVPR 2023poster

The ability of scale-equivariance processing blocks plays a central role in arbitrary-scale image super-resolution tasks. Inspired by this crucial observation, this work proposes two novel scale-equivariant modules within a transformer-style framework to enhance arbitrary-scale image super-resolutio…

2023

Learning Continuous Depth Representation via Geometric Spatial Aggregator

AAAI 2023technical

Depth map super-resolution (DSR) has been a fundamental task for 3D computer vision. While arbitrary scale DSR is a more realistic setting in this scenario, previous approaches predominantly suffer from the issue of inefficient real-numbered scale upsampling. To explicitly address this issue, we pro…

2023

Meta-Adapter: An Online Few-shot Learner for Vision-Language Model

NeurIPS 2023poster

The contrastive vision-language pre-training, known as CLIP, demonstrates remarkable potential in perceiving open-world visual concepts, enabling effective zero-shot image recognition. Nevertheless, few-shot learning methods based on CLIP typically require offline fine-tuning of the parameters on…

Cited by 13SourcePDFScholar
2023

Omni Aggregation Networks for Lightweight Image Super-Resolution

CVPR 2023poster

While lightweight ViT framework has made tremendous progress in image super-resolution, its uni-dimensional self-attention modeling, as well as homogeneous aggregation scheme, limit its effective receptive field (ERF) to include more comprehensive interactions from both spatial and channel dimension…

2023

Training Set Cleansing of Backdoor Poisoning by Self-Supervised Representation Learning

ICASSP 2023accepted

A backdoor or Trojan attack is an important type of data poisoning attack against deep neural network (DNN) classifiers, wherein the training dataset is poisoned with a small number of samples that each possess the backdoor pattern (usually a pattern that is either imperceptible or innocuous) and wh…

Cited by 0SourceScholar
2022

PAC-Net: Highlight Your Video via History Preference Modeling

ECCV 2022poster

"Autonomous highlight detection is crucial for video editing and video browsing on social media platforms. General video highlight detection aims at extracting the most interesting segments from the entire video. However, interest is subjective among different users. A naive solution is to train a m…

Cited by 4SourcePDFScholar
2021

3D Human Action Representation Learning via Cross-View Consistency Pursuit

CVPR 2021poster

In this work, we propose a Cross-view Contrastive Learning framework for unsupervised 3D skeleton-based action representation (CrosSCLR), by leveraging multi-view complementary supervision signal. CrosSCLR consists of both single-view contrastive learning (SkeletonCLR) and cross-view consistent know…

Cited by 240PDFcodeScholar
2021

Cross-Category Video Highlight Detection via Set-Based Learning

ICCV 2021poster

Autonomous highlight detection is crucial for enhancing the efficiency of video browsing on social media platforms. To attain this goal in a data-driven way, one may often face the situation where highlight annotations are not available on the target video category used in practice, while the superv…

Cited by 62PDFcodeScholar
2021

Self-supervised Graph-level Representation Learning with Local and Global Structure

ICML 2021spotlight

This paper studies unsupervised/self-supervised whole-graph representation learning, which is critical in many tasks such as molecule properties prediction in drug and material discovery. Existing methods mainly focus on preserving the local similarity structure between different graph instances but…

2021

Shape Self-Correction for Unsupervised Point Cloud Understanding

ICCV 2021poster

We develop a novel self-supervised learning method named Shape Self-Correction for point cloud analysis. Our method is motivated by the principle that a good shape representation should be able to find distorted parts of a shape and correct them. To learn strong shape representations in an unsupervi…

Cited by 58PDFScholar
2020

Cross-Domain Detection via Graph-Induced Prototype Alignment

CVPR 2020oral

Applying the knowledge of an object detector trained on a specific domain directly onto a new domain is risky, as the gap between two domains can severely degrade model's performance. Furthermore, since different instances commonly embody distinct modal information in object detection scenario, the…

Cited by 302PDFcodeScholar
2020

Learning to Combine: Knowledge Aggregation for Multi-Source Domain Adaptation

ECCV 2020poster

Transferring knowledges learned from multiple source domains to target domain is a more practical and challenging task than conventional single-source domain adaptation. Furthermore, the increase of modalities brings more difficulty in aligning feature distributions among multiple domains. To mitiga…