← Search

Dahyun Chung

4 accepted papers

2026

MATRIX: Mask Track Alignment for Interaction-aware Video Generation

ICLR 2026poster

Video DiTs have advanced video generation, yet they still struggle to model multi-instance or subject-object interactions. This raises a key question: How do these models internally represent interactions? To answer this, we curate MATRIX-11K, a video dataset with interaction-aware captions and mul…

Cited by 0SourcecodeScholar
2026

uCLIP: Parameter-Efficient Multilingual Extension of Vision-Language Models with Unpaired Data

AAAI 2026technical

Contrastive Language–Image Pre-training (CLIP) has demonstrated strong generalization across a wide range of visual tasks by leveraging large-scale English–image pairs. However, its extension to low-resource languages remains limited due to the scarcity of high-quality multilingual image–text data.

Cited by 0SourcePDFScholar
2025

AM-Adapter: Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-Wild

ICCV 2025poster

Exemplar-based semantic image synthesis generates images aligned with semantic content while preserving the appearance of an exemplar. Conventional structure-guidance models like ControlNet, are limited as they rely solely on text prompts to control appearance and cannot utilize exemplar images as i…

Cited by 0SourcePDFScholar
2025

Emergent Temporal Correspondences from Video Diffusion Transformers

NeurIPS 2025poster

Recent advancements in video diffusion models based on Diffusion Transformers (DiTs) have achieved remarkable success in generating temporally coherent videos. Yet, a fundamental question persists: how do these models internally establish and represent temporal correspondences across frames? We int…

Cited by 0SourceScholar