← Search

Yongsheng Yu

9 accepted papers

2026

PixelDiT: Pixel Diffusion Transformers for Image Generation

CVPR 2026

Latent-space modeling has been the standard for Diffusion Transformers (DiTs). However, it relies on a two-stage pipeline where the pretrained autoencoder introduces lossy reconstruction, leading to error accumulation while hindering joint optimization. To address these issues, we propose PixelDiT,

Cited by 0SourcecodeScholar
2026

RealUHR: Harnessing Patch-Cascade Flows for Photorealistic Ultra-High-Resolution Synthesis

AAAI 2026technical

Ultra-high-resolution (UHR) text-to-image synthesis faces significant hurdles, including immense computational costs and a scarcity of training data. To address these, we introduce RealUHR, an efficient and scalable framework for generating photorealistic 4K images. At its core, RealUHR employs a Pa

Cited by 0SourcePDFScholar
2025

OmniPaint: Mastering Object-Oriented Editing via Disentangled Insertion-Removal Inpainting

ICCV 2025poster

Diffusion-based generative models have revolutionized object-oriented image editing, yet their deployment in realistic object removal and insertion remains hampered by challenges such as the intricate interplay of physical effects and insufficient paired training data. In this work, we introduce Omn…

Cited by 0SourcePDFScholar
2025

PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement

NeurIPS 2025poster

Latent Diffusion Models (LDMs) have markedly advanced the quality of image inpainting and local editing. However, the inherent latent compression often introduces pixel-level inconsistencies, such as chromatic shifts, texture mismatches, and visible seams along editing boundaries. Existing remedies,…

Cited by 0SourceScholar
2023

Monocular 3D Human Pose Estimation Based on Global Temporal-Attentive and Joints-Attention In Video

ICASSP 2023accepted

Learning to capture human motion is essential to 3D human pose and shape estimation from monocular video, which is widely used in many 3D applications. However, the existing methods mainly rely on recurrent or convolutional operation to model such temporal information, which limits the ability to ca…

Cited by 0SourceScholar