← Search

Zili Yi

14 accepted papers

2026

A Training-Free Framework for High-Fidelity Appearance Transfer via Diffusion Transformers

ICASSP 2026poster

Diffusion Transformers (DiTs) excel at generation, but their global self-attention makes controllable, reference-image-based editing a distinct challenge. Unlike U-Nets, naively injecting local appearance into a DiT can disrupt its holistic scene structure. We address this by proposing the first tra…

Cited by 0SourcePDFScholar
2026

Listen and Count: Expanding the Frontier of Zero-Shot Object Counting to Sound-centric Counting

IJCAI 2026

While class-agnostic object counting has recently evolved from image-exemplar to language-guided paradigms, existing methods are limited by text polysemy and the lack of prompts in audio-sensing scenarios. To overcome these challenges, we introduce a sound-centric counting paradigm, enabling models

Cited by 0Scholar
2025

Anywhere: A Multi-Agent Framework for User-Guided, Reliable, and Diverse Foreground-Conditioned Image Generation

AAAI 2025technical

Recent advancements in image-conditioned image generation have demonstrated substantial progress. However, foreground-conditioned image generation remains underexplored, encountering challenges such as compromised object integrity, foreground-background inconsistencies, limited diversity, and reduce…

Cited by 0SourcePDFScholar
2025

FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction

NeurIPS 2025poster

Controlling the spatial and semantic structure of diffusion-generated images remains a challenge. Existing methods like ControlNet rely on handcrafted condition maps and retraining, limiting flexibility and generalization. Inversion-based approaches offer stronger alignment but incur high inference…

Cited by 0SourceScholar
2025

From Zero to Detail: Deconstructing Ultra-High-Definition Image Restoration from Progressive Spectral Perspective

CVPR 2025poster

Ultra-high-definition (UHD) image restoration faces significant challenges due to its high resolution, complex content, and intricate details. To cope with these challenges, we analyze the restoration process in depth through a progressive spectral perspective, and deconstruct the complex UHD restor…

2025

OmniStyle: Filtering High Quality Style Transfer Data at Scale

CVPR 2025poster

In this paper, we introduce OmniStyle-1M, a large-scale paired style transfer dataset comprising over one million content-style-stylized image triplets across 1,000 diverse style categories, each enhanced with textual descriptions and instruction prompts. We show that OmniStyle-1M can not only impro…

Cited by 0SourcePDFScholar
2025

One-Shot Learning for Pose-Guided Person Image Synthesis in the Wild

ICASSP 2025accepted

Current Pose-Guided Person Image Synthesis (PGPIS) methods depend heavily on large amounts of labeled triplet data to train the generator in a supervised manner. However, they often falter when applied to in-the-wild samples, primarily due to the distribution gap between the training datasets and re…

Cited by 4SourceScholar
2025

SigStyle: Signature Style Transfer via Personalized Text-to-Image Models

AAAI 2025technical

Style transfer enables the seamless integration of artistic styles from a style image into a content image, resulting in visually striking and aesthetically enriched outputs. Despite numerous advances in this field, existing methods did not explicitly focus on the signature style, which represents…

Cited by 1SourcePDFScholar
2024

SemanticHuman-HD: High Resolution Semantic disentangled 3D Human Generation

ECCV 2024poster

"With the development of neural radiance fields and generative models, numerous methods have been proposed for learning 3D human generation from 2D images. These methods allow control over the pose of the generated 3D human and enable rendering from different viewpoints. However, none of these metho…

2023

ReGANIE: Rectifying GAN Inversion Errors for Accurate Real Image Editing

AAAI 2023technical

The StyleGAN family succeed in high-fidelity image generation and allow for flexible and plausible editing of generated images by manipulating the semantic-rich latent style space. However, projecting a real image into its latent space encounters an inherent trade-off between inversion quality and e…

Cited by 7SourcePDFScholar
2022

XMP-Font: Self-Supervised Cross-Modality Pre-Training for Few-Shot Font Generation

CVPR 2022poster

Generating a new font library is a very labor-intensive and time-consuming job for glyph-rich scripts. Few-shot font generation is thus required, as it requires only a few glyph references without fine-tuning during test. Existing methods follow the style-content disentanglement paradigm, and expect…

Cited by 58PDFScholar
2020

Contextual Residual Aggregation for Ultra High-Resolution Image Inpainting

CVPR 2020oral

Recently data-driven image inpainting methods have made inspiring progress, impacting fundamental image editing tasks such as object removal and damaged image repairing. These methods are more effective than classic approaches, however, due to memory limitations they can only handle low-resolution i…

Cited by 443PDFcodeScholar