← Search

Zhenxiong Tan

11 accepted papers

2026

Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers

CVPR 2026

Recent advances in diffusion-based controllable visual generation have led to remarkable improvements in image quality. However, these powerful models are typically deployed on cloud servers due to their large computational demands, raising serious concerns about user data privacy. To enable secure

Cited by 0SourceScholar
2026

Minute-Long Videos with Dual Parallelisms

AAAI 2026technical

Diffusion Transformer (DiT)-based video diffusion models generate high-quality videos at scale but incur prohibitive processing latency and memory costs for long videos. To address this, we propose a novel distributed inference strategy, termed DualParal. The core idea is that, instead of generating

Cited by 0SourcePDFScholar
2026

SpotEdit: Selective Region Editing in Diffusion Transformers

CVPR 2026

Diffusion Transformer (DiT)-based models have significantly advanced image editing by encoding conditional images and integrating them into transformer layers. However, most edits involve modifying only small regions, while current methods uniformly process and denoise all tokens at every timestep,

Cited by 0SourcecodeScholar
2025

CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up

NeurIPS 2025poster

Diffusion Transformers (DiT) have become a leading architecture in image generation. However, the quadratic complexity of attention mechanisms, which are responsible for modeling token-wise relationships, results in significant latency when generating high-resolution images. To address this issue, w…

Cited by 0SourcecodeScholar
2025

Image Editing As Programs with Diffusion Models

NeurIPS 2025poster

While diffusion models have achieved remarkable success in text-to-image generation, they encounter significant challenges with instruction-driven image editing. Our research highlights a key challenge: these models particularly struggle with structurally-inconsistent edits that involve substantial…

Cited by 0SourcecodeScholar
2025

OminiControl: Minimal and Universal Control for Diffusion Transformer

ICCV 2025poster

We present OminiControl, a novel approach that rethinks how image conditions are integrated into Diffusion Transformer (DiT) architectures. Current image conditioning methods either introduce substantial parameter overhead or handle only specific control tasks effectively, limiting their practical v…

2024

AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising

NeurIPS 2024poster

Diffusion models have garnered significant interest from the community for their great generative ability across various applications. However, their typical multi-step sequential-denoising nature gives rise to high cumulative latency, thereby precluding the possibilities of parallel computation. To…

2024

MindBridge: A Cross-Subject Brain Decoding Framework

CVPR 2024highlight

Brain decoding a pivotal field in neuroscience aims to reconstruct stimuli from acquired brain signals primarily utilizing functional magnetic resonance imaging (fMRI). Currently brain decoding is confined to a per-subject-per-model paradigm limiting its applicability to the same individual for whom…

2020

AdversarialNAS: Adversarial Neural Architecture Search for GANs

CVPR 2020poster

Neural Architecture Search (NAS) that aims to automate the procedure of architecture design has achieved promising results in many computer vision fields. In this paper, we propose an AdversarialNAS method specially tailored for Generative Adversarial Networks (GANs) to search for a superior generat…

Cited by 114PDFcodeScholar