← Search

Seil Kang

7 accepted papers

2026

I'm a Map! Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers

CVPR 2026

Video Diffusion Transformers (DiTs) have been synthesizing high-quality video with high fidelity from given text descriptions involving motion. However, understanding how Video DiTs convert motion words into video remains insufficient. Furthermore, while prior studies on interpretable saliency maps

Cited by 0SourcecodeScholar
2026

Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

ICML 2026poster

Video diffusion models can generate visually stunning content, yet frequently produce motion that violates physical laws, objects accelerate implausibly or vanish mid-trajectory. We reveal a surprising finding: a 2-step generation often exhibits better physical consistency than a 50-step output from…

Cited by 0SourceScholar
2026

Real-Time Visual Attribution Streaming in Thinking Model

ICML 2026spotlight

We present an amortized framework for real-time visual attribution streaming in multimodal thinking models. When these models generate code from a screenshot or solve math problems from images, their long reasoning traces should be grounded in visual evidence. However, verifying this reliance is cha…

Cited by 0SourceScholar
2026

ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting

CVPR 2026

Recent advancements in Video Large Language Models (VideoLLMs) have enabled strong performance across diverse multimodal video tasks. To reduce the high computational cost of processing dense video frames, efficiency-oriented methods such as frame selection have been widely adopted. While effective

Cited by 0SourcecodeScholar
2025

Rare Text Semantics Were Always There in Your Diffusion Transformer

NeurIPS 2025poster

Starting from flow- and diffusion-based transformers, Multi-modal Diffusion Transformers (MM-DiTs) have reshaped text-to-vision generation, gaining acclaim for exceptional visual fidelity. As these models advance, users continually push the boundary with imaginative or rare prompts, which advanced m…

Cited by 0SourceScholar
2025

See What You Are Told: Visual Attention Sink in Large Multimodal Models

ICLR 2025poster

Large multimodal models (LMMs) "see" images by leveraging the attention mechanism between text and visual tokens in the transformer decoder. Ideally, these models should focus on key visual information relevant to the text token. However, recent findings indicate that LMMs have an extraordinary tend…

Cited by 2SourcePDFScholar
2025

Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding

CVPR 2025highlight

Visual grounding seeks to localize the image region corresponding to a free-form text description. Recently, the strong multimodal capabilities of Large Vision-Language Models (LVLMs) have driven substantial improvements in visual grounding, though they inevitably require fine-tuning and additional…

Cited by 2SourcePDFScholar