← Search

Haowei Zhu

8 accepted papers

2026

DIFFSPARSE: ACCELERATING DIFFUSION TRANSFORMERS WITH LEARNED TOKEN SPARSITY

ICLR 2026poster

Diffusion models demonstrate outstanding performance in image generation, but their multi-step inference mechanism requires immense computational cost. Previous works accelerate inference by leveraging layer or token cache techniques to reduce computational cost. However, these methods fail to achie…

Cited by 0SourceScholar
2026

DiffBench Meets DiffAgent: End-to-End LLM-Driven Diffusion Acceleration Code Generation

AAAI 2026technical

Diffusion models have achieved remarkable success in image and video generation. However, their inherently multiple step inference process imposes substantial computational overhead, hindering real-world deployment. Accelerating diffusion models is therefore essential, yet determining how to combine

Cited by 0SourcePDFScholar
2026

Mango-GS: Enhancing Spatio-Temporal Consistency in Dynamic Scenes Reconstruction using Multi-Frame Node-Guided 4D Gaussian Splatting

ICLR 2026poster

Reconstructing dynamic 3D scenes with photorealistic detail and temporal coherence remains a significant challenge. Existing Gaussian splatting approaches modeling scenes rely on per-frame optimization, causing them to overfit to instantaneous states rather than learning true motion dynamics. To add…

Cited by 0SourceScholar
2026

TAP: A Token-Adaptive Predictor Framework for Training-Free Diffusion Acceleration

CVPR 2026

Diffusion models achieve strong generative performance but remain slow at inference due to the need for repeated full-model denoising passes. We present Token-Adaptive Predictor (TAP), a training-free, probe-driven framework that adaptively selects a predictor for each token at every sampling step.

Cited by 0SourceScholar
2025

Accelerating Diffusion-based Super-Resolution with Dynamic Time-Spatial Sampling

IJCAI 2025

Diffusion models have gained attention for their success in modeling complex distributions, achieving impressive perceptual quality in SR tasks. However, existing diffusion-based SR methods often suffer from high computational costs, requiring numerous iterative steps for training and inference. Exi

Cited by 0SourcePDFScholar
2025

ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection

NeurIPS 2025spotlight

The scale and quality of datasets are crucial for training robust perception models. However, obtaining large-scale annotated data is both costly and time-consuming. Generative models have emerged as a powerful tool for data augmentation by synthesizing samples that adhere to desired distributions.…

Cited by 0SourceScholar
2025

Semantic Hierarchical Prompt Tuning for Parameter-Efficient Fine-Tuning

ICASSP 2025accepted

As the scale of vision models continues to grow, Visual Prompt Timing (VPT) has emerged as a parameter-efficient transfer learning technique, noted for its superior performance compared to full fine-tuning. However, indiscriminately applying prompts to every layer without considering their inherent…

Cited by 0SourceScholar
2022

Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-Identification

CVPR 2022poster

Recently, self-attention mechanisms have shown impressive performance in various NLP and CV tasks, which can help capture sequential characteristics and derive global information. In this work, we explore how to extend self-attention modules to better learn subtle feature embeddings for recognizing…

Cited by 230PDFScholar