← Search

Woojung Han

7 accepted papers

2026

I'm a Map! Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers

CVPR 2026

Video Diffusion Transformers (DiTs) have been synthesizing high-quality video with high fidelity from given text descriptions involving motion. However, understanding how Video DiTs convert motion words into video remains insufficient. Furthermore, while prior studies on interpretable saliency maps

Cited by 0SourcecodeScholar
2026

Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

ICML 2026poster

Video diffusion models can generate visually stunning content, yet frequently produce motion that violates physical laws, objects accelerate implausibly or vanish mid-trajectory. We reveal a surprising finding: a 2-step generation often exhibits better physical consistency than a 50-step output from…

Cited by 0SourceScholar
2026

Real-Time Visual Attribution Streaming in Thinking Model

ICML 2026spotlight

We present an amortized framework for real-time visual attribution streaming in multimodal thinking models. When these models generate code from a screenshot or solve math problems from images, their long reasoning traces should be grounded in visual evidence. However, verifying this reliance is cha…

Cited by 0SourceScholar
2025

Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation

CVPR 2025poster

Open-Vocabulary Semantic Segmentation (OVSS) has advanced with recent vision-language models (VLMs), enabling segmentation beyond predefined categories through various learning schemes. Notably, training-free methods offer scalable, easily deployable solutions for handling unseen data, a key goal of…

Cited by 1SourcePDFScholar
2025

Rare Text Semantics Were Always There in Your Diffusion Transformer

NeurIPS 2025poster

Starting from flow- and diffusion-based transformers, Multi-modal Diffusion Transformers (MM-DiTs) have reshaped text-to-vision generation, gaining acclaim for exceptional visual fidelity. As these models advance, users continually push the boundary with imaginative or rare prompts, which advanced m…

Cited by 0SourceScholar
2025

Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis

CVPR 2025poster

Diffusion-based text-to-image (T2I) models have recently excelled in high-quality image generation, particularly in a training-free manner, enabling cost-effective adaptability and generalization across diverse tasks. However, while the existing methods have been continuously focusing on several cha…

Cited by 0SourcePDFScholar
2024

EAGLE: Eigen Aggregation Learning for Object-Centric Unsupervised Semantic Segmentation

CVPR 2024highlight

Semantic segmentation has innately relied on extensive pixel-level annotated data leading to the emergence of unsupervised methodologies. Among them leveraging self-supervised Vision Transformers for unsupervised semantic segmentation (USS) has been making steady progress with expressive deep featur…