← Search

Bryan Sangwoo Kim

4 accepted papers

2025

Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment

NeurIPS 2025spotlight

Modern single-image super-resolution (SISR) models deliver photo-realistic results at the scale factors on which they are trained, but collapse when asked to magnify far beyond that regime. We address this scalability bottleneck with Chain-of-Zoom (CoZ), a model-agnostic framework that factorizes SI…

Cited by 0SourceScholar
2025

Free2Guide: Training-Free Text-to-Video Alignment using Image LVLM

ICCV 2025poster

Diffusion models have achieved impressive results in generative tasks for text-to-video (T2V) synthesis. However, achieving accurate text alignment in T2V generation remains challenging due to the complex temporal dependencies across frames. Existing reinforcement learning (RL)-based approaches to e…

2025

VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's Guide

CVPR 2025poster

Text-to-image (T2I) diffusion models have revolutionized visual content creation, but extending these capabilities to text-to-video (T2V) generation remains a challenge, particularly in preserving temporal consistency. Existing methods that aim to improve consistency often cause trade-offs such as r…