← Search

Hoigi Seo

9 accepted papers

2026

Erasing Thousands of Concepts: Towards Scalable and Practical Concept Erasure for Text-to-Image Diffusion Models

CVPR 2026

Large-scale text-to-image (T2I) diffusion models deliver remarkable visual fidelity but pose safety risks due to their capacity to reproduce undesirable content, such as copyrighted ones. Concept erasure has emerged as a mitigation strategy, yet existing approaches struggle to balance scalability, p

Cited by 0SourcecodeScholar
2026

Training-free Mixed-Resolution Latent Upsampling for Spatially Accelerated Diffusion Transformers

CVPR 2026

Diffusion transformers (DiTs) offer excellent scalability for high-fidelity generation, but their computational overhead poses a great challenge for practical deployment. Existing acceleration methods primarily exploit the temporal dimension, whereas spatial acceleration remains underexplored. In th

Cited by 0SourcecodeScholar
2026

Training-free, Perceptually Consistent Low-Resolution Previews with High-Resolution Image for Efficient Workflows of Diffusion Models

CVPR 2026

Image generative models have become indispensable tools to yield exquisite high-resolution (HR) images for everyone, ranging from general users to professional designers. However, a desired outcome often requires generating a large number of HR images with different prompts and seeds, resulting in h

Cited by 0SourceScholar
2026

Triadic Dynamics Aware Diffusion Posterior Sampling for Inverse Problems: Optimizing Guidance and Stochasticity Schedules

ICML 2026poster

Generative posterior sampling using diffusion models has emerged as a dominant paradigm for solving inverse problems in imaging, which usually consists of three main components: data consistency (DC) guidance, classifier-free guidance (CFG) and stochasticity. While prior arts have focused on how to …

Cited by 0SourceScholar
2025

Efficient Personalization of Quantized Diffusion Model without Backpropagation

CVPR 2025poster

Diffusion models have shown remarkable performance in image synthesis, but they demand extensive computational and memory resources for training, fine-tuning and inference. Although advanced quantization techniques have successfully minimized memory usage for inference, training and fine-tuning thes…

2025

On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models

NeurIPS 2025poster

Large vision-language models (LVLMs), which integrate a vision encoder (VE) with a large language model, have achieved remarkable success across various tasks. However, there are still crucial challenges in LVLMs such as object hallucination, generating descriptions of objects that are not in the in…

Cited by 0SourceScholar
2025

Skrr: Skip and Re-use Text Encoder Layers for Memory Efficient Text-to-Image Generation

ICML 2025poster

Large-scale text encoders in text-to-image (T2I) diffusion models have demonstrated exceptional performance in generating high-quality images from textual prompts. Unlike denoising modules that rely on multiple iterative steps, text encoders require only a single forward pass to produce text embeddi…

Cited by 0SourcePDFScholar
2024

BeyondScene: Higher-Resolution Human-Centric Scene Generation With Pretrained Diffusion

ECCV 2024poster

"Generating higher-resolution human-centric scenes with details and controls remains a challenge for existing text-to-image diffusion models. This challenge stems from limited training image size, text encoder capacity (limited tokens), and the inherent difficulty of generating complex scenes involv…

2024

INTRA: Interaction Relationship-aware Weakly Supervised Affordance Grounding

ECCV 2024poster

"Affordance denotes the potential interactions inherent in objects. The perception of affordance can enable intelligent agents to navigate and interact with new environments efficiently. Weakly supervised affordance grounding teaches agents the concept of affordance without costly pixel-level annota…

Cited by 4SourcePDFScholar