← Search

Joonghyuk Shin

3 accepted papers

2026

Direct Reward Fine-Tuning on Poses for Single Image to 3D Human in the Wild

ICLR 2026poster

Single-view 3D human reconstruction has achieved remarkable progress through the adoption of multi-view diffusion models, yet the recovered 3D humans often exhibit unnatural poses. This phenomenon becomes pronounced when reconstructing 3D humans with dynamic or challenging poses, which we attribute…

Cited by 0SourceScholar
2026

MotionStream: Real-Time Video Generation with Interactive Motion Controls

ICLR 2026oral

Current motion-conditioned video generation methods suffer from prohibitive latency (minutes per video) and non-causal processing that prevents real-time interaction. We present MotionStream, enabling sub-second latency with up to 29 FPS streaming generation on a single GPU. Our approach begins by a…

Cited by 0SourcecodeScholar
2025

Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing

ICCV 2025poster

Transformer-based diffusion models have recently superseded traditional U-Net architectures, with multimodal diffusion transformers (MM-DiT) emerging as the dominant approach in state-of-the-art models like Stable Diffusion 3 and Flux.1. Previous approaches have relied on unidirectional cross-attent…

Cited by 0SourcePDFScholar