← Search

Soo Ye Kim

25 accepted papers

2026

Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing

ICML 2026poster

Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To unify vision generation and understanding, a burgeoning trend is to adopt high-dimensional features from representation enc…

Cited by 0SourceScholar
2026

EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning

ICLR 2026oral

Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. While image generation and editing have rapidly transitioned from task-specific to unified frameworks, video generation and editing remain fragmented due…

Cited by 0SourcecodeScholar
2026

Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Model

ICLR 2026poster

Advancements in diffusion models have significantly improved video quality, directing attention to fine-grained controllability. However, many existing methods depend on fine-tuning large-scale video models for specific tasks, which becomes increasingly impractical as model sizes continue to grow. I…

Cited by 0SourcecodeScholar
2026

I-INR: Iterative Implicit Neural Representations

AAAI 2026technical

Implicit Neural Representations (INRs) have revolutionized signal processing and computer vision by modeling signals as continuous, differentiable functions parameterized by neural networks. However, INRs are prone to the spectral bias problem, limiting their ability to retain high-frequency informa

Cited by 0SourcePDFScholar
2026

LightMover: Generative Light Movement with Color and Intensity Controls

CVPR 2026

We present LightMover, a framework for controllable light manipulation in single images that leverages video diffusion priors to produce physically plausible illumination changes without re-rendering the scene. We formulate light editing as a sequence-to-sequence prediction problem in visual token s

Cited by 0SourceScholar
2026

OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation

ICLR 2026poster

Recent advancements in diffusion-based text synthesis have demonstrated significant performance in inserting and editing text within images via inpainting. However, despite the potential of text inpainting methods, three key limitations hinder their applicability to broader Text Image Manipulation (…

Cited by 0SourcecodeScholar
2026

PropFly: Learning to Propagate via On-the-Fly Supervision from Pre-trained Video Diffusion Models

CVPR 2026

Propagation-based video editing enables precise user control by propagating a single edited frame into following frames while maintaining the original context such as motion and structures.However, training such models requires large-scale, paired (source and edited) video datasets, which are costly

Cited by 0SourceScholar
2025

Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction

ICCV 2025poster

Existing feedforward image-to-3D methods mainly rely on 2D multi-view diffusion models that cannot guarantee 3D consistency. These methods easily collapse when changing the prompt view direction and mainly handle object-centric cases. In this paper, we propose a novel single-stage 3D diffusion model…

2025

FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity

CVPR 2025poster

The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate integration of visual and textual information across various applications, including image and video captioning, visual question answering, and cross-modal retrieva…

Cited by 6SourcePDFScholar
2025

Generative Image Layer Decomposition with Visual Effects

CVPR 2025poster

Recent advancements in large generative models, particularly diffusion-based methods, have significantly enhanced the capabilities of image editing. However, achieving precise control over image composition tasks remains a challenge. Layered representations, which allow for independent editing of im…

Cited by 1SourcePDFScholar
2025

Generative Video Propagation

CVPR 2025poster

Large-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative video propagation framework, various video tasks can be addressed in a unified way by leveraging the generative power of su…

Cited by 1SourcePDFScholar
2025

MetaShadow: Object-Centered Shadow Detection, Removal, and Synthesis

CVPR 2025poster

Shadows are often underconsidered or even ignored in image editing applications, limiting the realism of the edited results. In this paper, we introduce MetaShadow, a three-in-one versatile framework that enables detection, removal, and controllable synthesis of shadows in natural images in an objec…

Cited by 2SourcePDFScholar
2025

Multitwine: Multi-Object Compositing with Text and Layout Control

CVPR 2025highlight

We introduce the first generative model capable of simultaneous multi-object compositing, guided by both text and layout. Our model allows for the addition of multiple objects within a scene, capturing a range of interactions from simple positional relations (e.g., next to, in front of) to complex…

Cited by 1SourcePDFScholar
2025

ObjectMover: Generative Object Movement with Video Prior

CVPR 2025poster

Simple as it seems, moving an object to another location within an image is, in fact, a challenging image-editing task that requires re-harmonizing the lighting, adjusting the pose based on perspective, accurately filling occluded regions, and ensuring coherent synchronization of shadows and reflect…

Cited by 1SourcePDFScholar
2025

OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions

NeurIPS 2025poster

Existing feedforward subject-driven video customization methods mainly study single-subject scenarios due to the difficulty of constructing multi-subject training data pairs. Another challenging problem that how to use the signals such as depth, mask, camera, and text prompts to control and edit the…

Cited by 0SourcecodeScholar
2025

Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment

ICLR 2025poster

Personalized image generation has emerged from the recent advancements in generative models. However, these generated personalized images often suffer from localized artifacts such as incorrect logos, reducing fidelity and fine-grained identity details of the generated results. Furthermore, there is…

Cited by 1SourcePDFScholar
2025

UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

CVPR 2025highlight

We introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs and outputs while capturing visual variations. Inspired by recent video generation…

2024

Descanning: From Scanned to the Original Images with a Color Correction Diffusion Model

AAAI 2024technical

A significant volume of analog information, i.e., documents and images, have been digitized in the form of scanned copies for storing, sharing, and/or analyzing in the digital world. However, the quality of such contents is severely degraded by various distortions caused by printing, storing, and sc…

Cited by 2SourcePDFScholar
2024

IMPRINT: Generative Object Compositing by Learning Identity-Preserving Representation

CVPR 2024poster

Generative object compositing emerges as a promising new avenue for compositional image editing. However the requirement of object identity preservation poses a significant challenge limiting practical usage of most existing methods. In response this paper introduces IMPRINT a novel diffusion-based…

Cited by 29SourcePDFScholar
2024

Thinking Outside the BBox: Unconstrained Generative Object Compositing

ECCV 2024poster

"Compositing an object into an image involves multiple non-trivial sub-tasks such as object placement and scaling, color/lighting harmonization, viewpoint/geometry adjustment, and shadow/reflection generation. Recent generative image compositing methods leverage diffusion models to handle multiple s…

Cited by 9SourcePDFScholar
2023

Modernizing Old Photos Using Multiple References via Photorealistic Style Transfer

CVPR 2023poster

This paper firstly presents old photo modernization using multiple references by performing stylization and enhancement in a unified manner. In order to modernize old photos, we propose a novel multi-reference-based old photo modernization (MROPM) framework consisting of a network MROPM-Net and a no…

Cited by 15SourcePDFScholar
2023

ObjectStitch: Object Compositing With Diffusion Model

CVPR 2023poster

Object compositing based on 2D images is a challenging problem since it typically involves multiple processing stages such as color harmonization, geometry correction and shadow generation to generate realistic results. Furthermore, annotating training data pairs for compositing requires substantial…

Cited by 94SourcePDFScholar
2021

KOALAnet: Blind Super-Resolution Using Kernel-Oriented Adaptive Local Adjustment

CVPR 2021poster

Blind super-resolution (SR) methods aim to generate a high quality high resolution image from a low resolution image containing unknown degradations. However, natural images contain various types and amounts of blur: some may be due to the inherent degradation characteristics of the camera, but some…

Cited by 92PDFcodeScholar
2019

Deep SR-ITM: Joint Learning of Super-Resolution and Inverse Tone-Mapping for 4K UHD HDR Applications

ICCV 2019oral

Recent modern displays are now able to render high dynamic range (HDR), high resolution (HR) videos of up to 8K UHD (Ultra High Definition). Consequently, UHD HDR broadcasting and streaming have emerged as high quality premium services. However, due to the lack of original UHD HDR video content, app…

Cited by 125PDFcodeScholar