← Search

Huijie Zhang

10 accepted papers

2026

AlphaFlow: Understanding and Improving MeanFlow Models

ICLR 2026poster

MeanFlow has recently emerged as a powerful framework for few-step generative modeling trained from scratch, but its success is not yet fully understood. In this work, we show that the MeanFlow objective naturally decomposes into two parts: trajectory flow matching and trajectory consistency. Throug…

Cited by 0SourcecodeScholar
2026

Quota-Calibrated Fine-Grained Alignment with Context-Aware Marginals for Text-based Person Retrieval

CVPR 2026

The core challenge in Text-based Person Retrieval (TPR) lies in establishing fine-grained, many-to-many semantic alignment between textual words and visual regions. Existing methods predominantly rely on pointwise similarity or attention mechanisms, implicitly assuming matches are independent and ba

Cited by 0SourceScholar
2025

A Closer Look at Model Collapse: From a Generalization-to-Memorization Perspective

NeurIPS 2025spotlight

The widespread use of diffusion models has led to an abundance of AI-generated data, raising concerns about model collapse---a phenomenon in which recursive iterations of training on synthetic data lead to performance degradation. Prior work primarily characterizes this collapse via variance shrinka…

Cited by 0SourceScholar
2025

An Efficient Pore Annotation Framework for Tight Sandstone Images with Segment Anything Model

ICASSP 2025accepted

Analyzing pore structures in tight sandstone thin sections is pivotal for assessing reservoir quality and predicting hydrocarbon migration, providing critical guidance for the formulation of energy extraction strategies. However, the significant variability in pore morphology, scale, and quantity re…

Cited by 0SourceScholar
2025

Shallow Diffuse: Robust and Invisible Watermarking through Low-Dim Subspaces in Diffusion Models

NeurIPS 2025spotlight

The widespread use of AI-generated content from diffusion models has raised significant concerns regarding misinformation and copyright infringement. Watermarking is a crucial technique for identifying these AI-generated images and preventing their misuse. In this paper, we introduce *Shallow Diffus…

Cited by 0SourceScholar
2024

Exploring Low-Dimensional Subspace in Diffusion Models for Controllable Image Editing

NeurIPS 2024poster

Recently, diffusion models have emerged as a powerful class of generative models. Despite their success, there is still limited understanding of their semantic spaces. This makes it challenging to achieve precise and disentangled image generation without additional training, especially in an unsupe…

2024

Improving Training Efficiency of Diffusion Models via Multi-Stage Framework and Tailored Multi-Decoder Architecture

CVPR 2024poster

Diffusion models emerging as powerful deep generative tools excel in various applications. They operate through a two-steps process: introducing noise into training samples and then employing a model to convert random noise into new samples (e.g. images). However their remarkable generative performa…

Cited by 12SourcePDFScholar
2024

The Emergence of Reproducibility and Consistency in Diffusion Models

ICML 2024poster

In this work, we investigate an intriguing and prevalent phenomenon of diffusion models which we term as "consistent model reproducibility'': given the same starting noise input and a deterministic sampler, different diffusion models often yield remarkably similar outputs. We confirm this phenomenon…

Cited by 57SourcePDFScholar
2022

ClearPose: Large-Scale Transparent Object Dataset and Benchmark

ECCV 2022poster

"Transparent objects are ubiquitous in household settings and pose distinct challenges for visual sensing and perception systems. The optical properties of transparent objects leaves conventional 3D sensors alone unreliable for object depth and pose estimation. These challenges are highlighted by th…

2022

ProgressLabeller: Visual Data Stream Annotation for Training Object-Centric 3D Perception

IROS 2022poster

Visual perception tasks often require vast amounts of labelled data, including 3D poses and image space segmen-tation masks. The process of creating such training data sets can prove difficult or time-intensive to scale up to efficacy for general use. Consider the task of pose estimation for rigid o…

Cited by 9SourcecodeScholar