← Search

Siyoon Jin

6 accepted papers

2026

MATRIX: Mask Track Alignment for Interaction-aware Video Generation

ICLR 2026poster

Video DiTs have advanced video generation, yet they still struggle to model multi-instance or subject-object interactions. This raises a key question: How do these models internally represent interactions? To answer this, we curate MATRIX-11K, a video dataset with interaction-aware captions and mul…

Cited by 0SourcecodeScholar
2026

Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry

AAAI 2026technical

We introduce a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view video data for training. Traditional reconstruction methods struggle with

Cited by 0SourcePDFScholar
2025

AM-Adapter: Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-Wild

ICCV 2025poster

Exemplar-based semantic image synthesis generates images aligned with semantic content while preserving the appearance of an exemplar. Conventional structure-guidance models like ControlNet, are limited as they rely solely on text prompts to control appearance and cannot utilize exemplar images as i…

Cited by 0SourcePDFScholar
2025

Emergent Temporal Correspondences from Video Diffusion Transformers

NeurIPS 2025poster

Recent advancements in video diffusion models based on Diffusion Transformers (DiTs) have achieved remarkable success in generating temporally coherent videos. Yet, a fundamental question persists: how do these models internally establish and represent temporal correspondences across frames? We int…

Cited by 0SourceScholar
2025

MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation

AAAI 2025technical

Conventional GAN-based models for talking head generation often suffer from limited quality and unstable training. Recent approaches based on diffusion models have attempted to address these limitations and improve fidelity. However, they still face challenges, such as intensive sampling times and d…

2024

DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image Personalization

CVPR 2024poster

The objective of text-to-image (T2I) personalization is to customize a diffusion model to a user-provided reference concept generating diverse images of the concept aligned with the target prompts. Conventional methods representing the reference concepts using unique text embeddings often fail to ac…