← Search

Zhongdao Wang

32 accepted papers

2026

Grounding Everything in Tokens for Multimodal Large Language Models

CVPR 2026

Multimodal large language models (MLLMs) have made significant advancements in vision understanding and reasoning. However, the autoregressive Transformer architecture used by MLLMs requires tokenization on input images, which limits their ability to accurately ground objects within the 2D image spa

Cited by 0SourceScholar
2025

Bi-Stream Knowledge Transfer for Semi-Supervised 3D Point Cloud Object Detection

ICRA 2025

3D point cloud object detection plays an important role in autonomous driving. However, labeling 3D object boxes is expensive and time-consuming, limiting the number of annotated point clouds used in fully-supervised training. This has led to a rise in semi-supervised 3D object detection research, w

Cited by 0SourceScholar
2025

Efficient 3D Perception on Multi-Sweep Point Cloud with Gumbel Spatial Pruning

ICRA 2025

This paper studies point cloud perception within outdoor environments. Existing methods face limitations in recognizing objects located at a distance or occluded, due to the sparse nature of outdoor point clouds. In this work, we observe a significant mitigation of this problem by accumulating multi

Cited by 1SourceScholar
2025

Enhancing LLM Knowledge Learning through Generalization

EMNLP 2025

As Large language models (LLMs) are increasingly deployed in diverse applications, faithfully integrating evolving factual knowledge into these models remains a critical challenge. Continued pre-training on paraphrased data has shown empirical promise for enhancing knowledge acquisition. However, th

2025

JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent

NeurIPS 2025poster

Photo retouching has become integral to contemporary visual storytelling, enabling users to capture aesthetics and express creativity. While professional tools such as Adobe Lightroom offer powerful capabilities, they demand substantial expertise and manual effort. In contrast, existing AI-based sol…

Cited by 0SourceScholar
2025

PVMamba: Parallelizing Vision Mamba via Dynamic State Aggregation

ICCV 2025poster

Mamba, an architecture with RNN-like sequence modeling of State Space Model (SSM), has demonstrated promising capabilities in long-range modeling with high efficiency. However, Mamba models struggle with structured 2D visual data using sequential computing, thereby lagging behind their attention-bas…

2025

ReNeg: Learning Negative Embedding with Reward Guidance

CVPR 2025highlight

In text-to-image (T2I) generation applications, negative embeddings have proven to be a simple yet effective approach for enhancing generation quality. Typically, these negative embeddings are derived from user-defined negative prompts, which, while being functional, are not necessarily optimal. In…

2025

Reliable and Calibrated Semantic Occupancy Prediction by Hybrid Uncertainty Learning

IJCAI 2025

Vision-centric semantic occupancy prediction plays a crucial role in autonomous driving, which requires accurate and reliable predictions from low-cost sensors. Although having notably narrowed the accuracy gap with LiDAR, there is still few research effort to explore the reliability and calibration

Cited by 0SourcePDFScholar
2025

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization

ICML 2025poster

Recent advancements in reinforcement learning from human feedback have shown that utilizing fine-grained token-level reward models can substantially enhance the performance of Proximal Policy Optimization (PPO) in aligning large language models. However, it is challenging to leverage such token-leve…

2025

Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis

ICCV 2025poster

Demand for 2K video synthesis is rising with increasing consumer expectations for ultra-clear visuals.While diffusion transformers (DiTs) have demonstrated remarkable capabilities in high-quality video generation, scaling them to 2K resolution remains computationally prohibitive due to quadratic gro…

Cited by 0SourcePDFScholar
2025

TurboVSR: Fantastic Video Upscalers and Where to Find Them

ICCV 2025poster

Diffusion-based generative models have demonstrated exceptional promise in the video super-resolution (VSR) task, achieving a substantial advancement in detail generation relative to prior methods. However, these approaches face significant computational efficiency challenges. For instance, current…

Cited by 0SourcePDFScholar
2024

LogoSticker: Inserting Logos into Diffusion Models for Customized Generation

ECCV 2024poster

"Recent advances in text-to-image model customization have underscored the importance of integrating new concepts with a few examples. Yet, these progresses are largely confined to widely recognized subjects, which can be learned with relative ease through models’ adequate shared prior knowledge. In…

2024

OccGen: Generative Multi-modal 3D Occupancy Prediction for Autonomous Driving

ECCV 2024poster

"Existing 3D semantic occupancy prediction methods typically treat the task as a one-shot 3D voxel-wise segmentation problem, focusing on a single-step mapping between the inputs and occupancy maps, which limits their ability to refine and complete local regions gradually. In this paper, we introduc…

Cited by 21SourcePDFScholar
2024

PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

ICLR 2024spotlight

The most advanced text-to-image (T2I) models require significant training costs (e.g., millions of GPU hours), seriously hindering the fundamental innovation for the AIGC community while increasing CO2 emissions. This paper introduces PixArt-$\alpha$, a Transformer-based T2I diffusion model whose im…

Cited by 546SourcePDFScholar
2024

PixArt-Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

ECCV 2024poster

"In this paper, we introduce , a Diffusion Transformer model (DiT) capable of directly generating images at 4K resolution. represents a significant advancement over its predecessor, , offering images of markedly higher fidelity and improved alignment with text prompts. A key feature of is its traini…

Cited by 178SourcePDFScholar
2024

QuadMamba: Learning Quadtree-based Selective Scan for Visual State Space Model

NeurIPS 2024poster

Recent advancements in State Space Models, notably Mamba, have demonstrated superior performance over the dominant Transformer models, particularly in reducing the computational complexity from quadratic to linear. Yet, difficulties in adapting Mamba from language to vision tasks arise due to the di…

2024

SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction

CVPR 2024poster

Vision-based perception for autonomous driving requires an explicit modeling of a 3D space where 2D latent representations are mapped and subsequent 3D operators are applied. However operating on dense latent spaces introduces a cubic time and space complexity which limits scalability in terms of pe…

Cited by 42SourcePDFScholar
2024

Taming Diffusion Prior for Image Super-Resolution with Domain Shift SDEs

NeurIPS 2024poster

Diffusion-based image super-resolution (SR) models have attracted substantial interest due to their powerful image restoration capabilities. However, prevailing diffusion models often struggle to strike an optimal balance between efficiency and performance. Typically, they either neglect to exploit…

2024

VEON: Vocabulary-Enhanced Occupancy Prediction

ECCV 2024poster

"Perceiving the world as 3D occupancy supports embodied agents to avoid collision with any types of obstacle. While open-vocabulary image understanding has prospered recently, how to bind the predicted 3D occupancy grids with open-world semantics still remains under-explored due to limited open-worl…

2023

Identity-Seeking Self-Supervised Representation Learning for Generalizable Person Re-Identification

ICCV 2023oral

This paper aims to learn a domain-generalizable (DG) person re-identification (ReID) representation from large-scale videos without any annotation. Prior DG ReID methods employ limited labeled data for training due to the high cost of annotation, which restricts further advances. To overcome the bar…

Cited by 28PDFcodeScholar
2023

MetaBEV: Solving Sensor Failures for 3D Detection and Map Segmentation

ICCV 2023poster

Perception systems in modern autonomous driving vehicles typically take inputs from complementary multi-modal sensors, e.g., LiDAR and cameras. However, in real-world applications, sensor corruptions and failures lead to inferior performances, thus compromising autonomous safety. In this paper, we p…

Cited by 43PDFScholar
2022

How to Synthesize a Large-Scale and Trainable Micro-Expression Dataset?

ECCV 2022poster

"This paper does not contain technical novelty but introduces our key discoveries in a data generation protocol, a database and insights. We aim to address the lack of large-scale datasets in micro-expression (MiE) recognition due to the prohibitive cost of data collection, which renders large-scale…

2022

Progressive-Granularity Retrieval Via Hierarchical Feature Alignment for Person Re-Identification

ICASSP 2022accepted

Person re-identification (re-ID) aims to match pedestrian images from non-overlapping cameras. It is a challenging task because of the feature misalignment problem caused by occlusion. In this paper, inspired by the coarse-to-fine nature of human perception, we propose a novel Progressive-Granularit…

Cited by 0SourceScholar
2022

Reliability-Aware Prediction via Uncertainty Learning for Person Image Retrieval

ECCV 2022poster

"Current person image retrieval methods have achieved great improvements in accuracy metrics. However, they rarely describe the reliability of the prediction. In this paper, we propose an Uncertainty-Aware Learning (UAL) method to remedy this issue. UAL aims at providing reliability-aware prediction…

2021

Do Different Tracking Tasks Require Different Appearance Models?

NeurIPS 2021poster

Tracking objects of interest in a video is one of the most popular and widely applicable problems in computer vision. However, with the years, a Cambrian explosion of use cases and benchmarks has fragmented the problem in a multitude of different experimental setups. As a consequence, the literature…

2020

Circle Loss: A Unified Perspective of Pair Similarity Optimization

CVPR 2020oral

This paper provides a pair similarity optimization viewpoint on deep feature learning, aiming to maximize the within-class similarity s_p and minimize the between-class similarity s_n. We find a majority of loss functions, including the triplet loss and the softmax cross-entropy loss, embed s_n and…

Cited by 1174PDFScholar
2020

CycAs: Self-supervised Cycle Association for Learning Re-identifiable Descriptions

ECCV 2020poster

This paper proposes a self-supervised learning method for the person re-identification (re-ID) problem, where existing unsupervised methods usually rely on pseudo labels, such as those from video tracklets or clustering. A potential drawback of using pseudo labels is that errors may accumulate and i…

Cited by 116SourcePDFScholar
2017

Orientation Invariant Feature Embedding and Spatial Temporal Regularization for Vehicle Re-Identification

ICCV 2017poster

In this paper, we tackle the vehicle Re-identification (ReID) problem which is of great importance in urban surveillance and can be used for multiple applications. In our vehicle ReID framework, an orientation invariant feature embedding module and a spatial-temporal regularization module are propos…

Cited by 458PDFScholar