← Search

Yuanzhi Zhu

11 accepted papers

2026

One-Step Flow for Image Super-Resolution with Tunable Fidelity-Realism Trade-offs

ICLR 2026poster

Recent advances in diffusion and flow-based generative models have demonstrated remarkable success in image restoration tasks, achieving superior perceptual quality compared to traditional deep learning approaches. However, these methods either require numerous sampling steps to generate high-qualit…

Cited by 0SourcecodeScholar
2025

Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation

ICCV 2025poster

Autoregressive visual generation models typically rely on tokenizers to compress images into tokens that can be predicted sequentially. A fundamental dilemma exists in token representation: discrete tokens enable straightforward modeling with standard cross-entropy loss, but suffer from information…

2025

Di[M]O: Distilling Masked Diffusion Models into One-step Generator

ICCV 2025poster

Masked Diffusion Models (MDMs) have emerged as a powerful generative modeling technique. Despite their remarkable results, they typically suffer from slow inference with several steps. In this paper, we propose Di\mathtt [M] O, a novel approach that distills masked diffusion models into a one-step g…

2025

Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to Advances

ICLR 2025poster

Current image watermarking methods are vulnerable to advanced image editing techniques enabled by large-scale text-to-image models. These models can distort embedded watermarks during editing, posing significant challenges to copyright protection. In this work, we introduce W-Bench, the first compre…

2024

Visual Text Generation in the Wild

ECCV 2024poster

"Recently, with the rapid advancements of generative models, the field of visual text generation has witnessed significant progress. However, it is still challenging to render high-quality text images in real-world scenarios, as three critical criteria should be satisfied: (1) Fidelity: the generate…

2023

Conditional Text Image Generation With Diffusion Models

CVPR 2023poster

Current text recognition systems, including those for handwritten scripts and scene text, have relied heavily on image synthesis and augmentation, since it is difficult to realize real-world complexity and diversity through collecting and annotating enough real text images. In this paper, we explore…

2023

DDFM: Denoising Diffusion Model for Multi-Modality Image Fusion

ICCV 2023oral

Multi-modality image fusion aims to combine different modalities to produce fused images that retain the complementary features of each modality, such as functional highlights and texture details. To leverage strong generative priors and address challenges such as unstable training and lack of inter…

Cited by 210PDFcodeScholar
2021

Implicit Feature Alignment: Learn To Convert Text Recognizer to Text Spotter

CVPR 2021poster

Text recognition is a popular research subject with many associated challenges. Despite the considerable progress made in recent years, the text recognition task itself is still constrained to solve the problem of reading cropped line text images and serves as a subtask of optical character recognit…

Cited by 16PDFcodeScholar
2020

Learn to Augment: Joint Data Augmentation and Network Optimization for Text Recognition

CVPR 2020poster

Handwritten text and scene text suffer from various shapes and distorted patterns. Thus training a robust recognition model requires a large amount of data to cover diversity as much as possible. In contrast to data collection and annotation, data augmentation is a low cost way. In this paper, we pr…

Cited by 116PDFcodeScholar
2019

Aggregation Cross-Entropy for Sequence Recognition

CVPR 2019oral

In this paper, we propose a novel method, aggregation cross-entropy (ACE), for sequence recognition from a brand new perspective. The ACE loss function exhibits competitive performance to CTC and the attention mechanism, with much quicker implementation (as it involves only four fundamental formulas…

Cited by 141PDFcodeScholar