← Search

Gwanhyeong Koo

12 accepted papers

2026

GADA: Geometry-Aware Deformable Aggregation for Image-Based Gaussian Splatting

ICML 2026poster

Gaussian Splatting has achieved significant improvements by incorporating warping-based techniques. These approaches enhance synthesis quality by warping images from source views into the target viewpoint to compensate for missing or residual pixels. However, such methods suffer from pixel-level ina…

Cited by 0SourceScholar
2026

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning

CVPR 2026

Reinforcement Learning with Verifiable Rewards (RLVR) traditionally relies on a sparse, outcome-based signal. Recent work shows that providing a fine-grained, model-intrinsic signal--rewarding the confidence growth in the ground-truth answer--effectively improves language reasoning training by provi

Cited by 0SourcecodeScholar
2025

FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow Fields

ICML 2025spotlight

Drag-based editing allows precise object manipulation through point-based control, offering user convenience. However, current methods often suffer from a geometric inconsistency problem by focusing exclusively on matching user-defined points, neglecting the broader geometry and leading to artifacts…

Cited by 0SourcePDFScholar
2025

ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On

CVPR 2025poster

This paper introduces ITA-MDT, the Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On (IVTON), designed to overcome the limitations of previous approaches by leveraging the Masked Diffusion Transformer (MDT) for improved handling of both global garment cont…

Cited by 0SourcePDFScholar
2025

Occlusion-robust Stylization for Drawing-based 3D Animation

ICCV 2025poster

3D animation aims to generate a 3D animated video from an input image and a target 3D motion sequence. Recent advances in image-to-3D models enable the creation of animations directly from user-hand drawings. Distinguished from conventional 3D animation, drawing-based 3D animation is crucial to pres…

Cited by 0SourcePDFScholar
2024

DNI: Dilutional Noise Initialization for Diffusion Video Editing

ECCV 2024poster

"Text-based diffusion video editing systems have been successful in performing edits with high fidelity and textual alignment. However, this success is limited to rigid-type editing such as style transfer and object overlay, while preserving the original structure of the input video. This limitation…

Cited by 2SourcePDFScholar
2024

FRAG: Frequency Adapting Group for Diffusion Video Editing

ICML 2024poster

In video editing, the hallmark of a quality edit lies in its consistent and unobtrusive adjustment. Modification, when integrated, must be smooth and subtle, preserving the natural flow and aligning seamlessly with the original vision. Therefore, our primary focus is on overcoming the current challe…

2024

FlexiEdit: Frequency-Aware Latent Refinement for Enhanced Non-Rigid Editing

ECCV 2024poster

"Current image editing methods primarily utilize DDIM Inversion, employing a two-branch diffusion approach to preserve the attributes and layout of the original image. However, these methods encounter challenges with non-rigid edits, which involve altering the image’s layout or structure. Our compre…

2024

Query-based Cross-Modal Projector Bolstering Mamba Multimodal LLM

EMNLP 2024finding

The Transformer’s quadratic complexity with input length imposes an unsustainable computational load on large language models (LLMs). In contrast, the Selective Scan Structured State-Space Model, or Mamba, addresses this computational challenge effectively. This paper explores a query-based cross-mo…

Cited by 0SourcePDFScholar
2024

TPC: Test-time Procrustes Calibration for Diffusion-based Human Image Animation

NeurIPS 2024poster

Human image animation aims to generate a human motion video from the inputs of a reference human image and a target motion video. Current diffusion-based image animation systems exhibit high precision in transferring human identity into targeted motion, yet they still exhibit irregular quality in th…

Cited by 3SourcePDFScholar
2024

Wavelet-Guided Acceleration of Text Inversion in Diffusion-Based Image Editing

ICASSP 2024accepted

In the field of image editing, Null-text Inversion (NTI) enables fine-grained editing while preserving the structure of the original image by optimizing null embeddings during the DDIM sampling process. However, the NTI process is time-consuming, taking more than two minutes per image. To address th…

Cited by 0SourceScholar
2023

SCANet: Scene Complexity Aware Network for Weakly-Supervised Video Moment Retrieval

ICCV 2023poster

Video moment retrieval aims to localize moments in video corresponding to a given language query. To avoid the expensive cost of annotating the temporal moments, weakly-supervised VMR (wsVMR) systems have been studied. For such systems, generating a number of proposals as moment candidates and then…

Cited by 21PDFScholar