← Search

Zhaoqing Wang

10 accepted papers

2026

When Safety Collides: Resolving Multi-Category Harmful Conflicts in Text-to-Image Diffusion via Adaptive Safety Guidance

CVPR 2026

Text-to-Image (T2I) diffusion models have demonstrated significant advancements in generating high-quality images, while raising potential safety concerns regarding harmful content generation. Safety-guidance-based methods have been proposed to mitigate harmful outputs by steering generation away fr

Cited by 0SourcecodeScholar
2025

Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video Generation

NeurIPS 2025poster

Text-to-Audio-Video (T2AV) generation aims to produce temporally and semantically aligned visual and auditory content from natural language descriptions. While recent progress in text-to-audio and text-to-video models has improved generation quality within each modality, jointly modeling them remain…

Cited by 0SourceScholar
2025

LaVin-DiT: Large Vision Diffusion Transformer

CVPR 2025poster

This paper presents the Large Vision Diffusion Transformer (LaVin-DiT), a scalable and unified foundation model designed to tackle over 20 computer vision tasks in a generative framework. Unlike existing large vision models directly adapted from natural language processing architectures, which rely…

2024

IDEAL: Influence-Driven Selective Annotations Empower In-Context Learners in Large Language Models

ICLR 2024poster

In-context learning is a promising paradigm that utilizes in-context examples as prompts for the predictions of large language models. These prompts are crucial for achieving strong performance. However, since the prompts need to be sampled from a large volume of annotated examples, finding the righ…

Cited by 28SourcePDFScholar
2023

BEV-SAN: Accurate BEV 3D Object Detection via Slice Attention Networks

CVPR 2023poster

Bird's-Eye-View (BEV) 3D Object Detection is a crucial multi-view technique for autonomous driving systems. Recently, plenty of works are proposed, following a similar paradigm consisting of three essential components, i.e., camera feature extraction, BEV feature construction, and task heads. Among…

Cited by 28SourcePDFScholar
2023

Mosaic Representation Learning for Self-supervised Visual Pre-training

ICLR 2023top-25%

Self-supervised learning has achieved significant success in learning visual representations without the need for manual annotation. To obtain generalizable representations, a meticulously designed data augmentation strategy is one of the most crucial parts. Recently, multi-crop strategies utilizing…

2022

CRIS: CLIP-Driven Referring Image Segmentation

CVPR 2022poster

Referring image segmentation aims to segment a referent via a natural linguistic expression. Due to the distinct data properties between text and image, it is challenging for a network to well align text and pixel-level features. Existing approaches use pretrained models to facilitate learning, yet…

Cited by 441PDFcodeScholar
2022

Exploring Set Similarity for Dense Self-Supervised Representation Learning

CVPR 2022poster

By considering the spatial correspondence, dense self-supervised representation learning has achieved superior performance on various dense prediction tasks. However, the pixel-level correspondence tends to be noisy because of many similar misleading pixels, e.g., backgrounds. To address this issue,…

Cited by 51PDFcodeScholar
2022

RSA: Reducing Semantic Shift from Aggressive Augmentations for Self-supervised Learning

NeurIPS 2022accept

Most recent self-supervised learning methods learn visual representation by contrasting different augmented views of images. Compared with supervised learning, more aggressive augmentations have been introduced to further improve the diversity of training pairs. However, aggressive augmentations may…

2021

Overfitting the Data: Compact Neural Video Delivery via Content-Aware Feature Modulation

ICCV 2021poster

Internet video delivery has undergone a tremendous explosion of growth over the past few years. However, the quality of video delivery system greatly depends on the Internet bandwidth. Deep Neural Networks (DNNs) are utilized to improve the quality of video delivery recently. These methods divide a…

Cited by 38PDFcodeScholar