← Search

Jisu Nam

10 accepted papers

2026

Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression

ICML 2026poster

Recent advances in autoregressive video diffusion have enabled real-time frame streaming, yet existing solutions still suffer from temporal repetition, drift, and motion deceleration. We find that naïvely applying StreamingLLM-style attention sinks to video diffusion leads to fidelity degradation an…

Cited by 0SourceScholar
2026

MATRIX: Mask Track Alignment for Interaction-aware Video Generation

ICLR 2026poster

Video DiTs have advanced video generation, yet they still struggle to model multi-instance or subject-object interactions. This raises a key question: How do these models internally represent interactions? To answer this, we curate MATRIX-11K, a video dataset with interaction-aware captions and mul…

Cited by 0SourcecodeScholar
2025

AM-Adapter: Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-Wild

ICCV 2025poster

Exemplar-based semantic image synthesis generates images aligned with semantic content while preserving the appearance of an exemplar. Conventional structure-guidance models like ControlNet, are limited as they rely solely on text prompts to control appearance and cannot utilize exemplar images as i…

Cited by 0SourcePDFScholar
2025

Emergent Temporal Correspondences from Video Diffusion Transformers

NeurIPS 2025poster

Recent advancements in video diffusion models based on Diffusion Transformers (DiTs) have achieved remarkable success in generating temporally coherent videos. Yet, a fundamental question persists: how do these models internally establish and represent temporal correspondences across frames? We int…

Cited by 0SourceScholar
2025

MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation

AAAI 2025technical

Conventional GAN-based models for talking head generation often suffer from limited quality and unstable training. Recent approaches based on diffusion models have attempted to address these limitations and improve fidelity. However, they still face challenges, such as intensive sampling times and d…

2025

Visual Persona: Foundation Model for Full-Body Human Customization

CVPR 2025poster

We introduce Visual Persona, a foundation model for text-to-image full-body human customization that, given a single in-the-wild human image, generates diverse images of the individual guided by text descriptions. Unlike prior methods that focus solely on preserving facial identity, our approach cap…

Cited by 0SourcePDFScholar
2024

Diffusion Model for Dense Matching

ICLR 2024oral

The objective for establishing dense correspondence between paired images con- sists of two terms: a data term and a prior term. While conventional techniques focused on defining hand-designed prior terms, which are difficult to formulate, re- cent approaches have focused on learning the data term w…

2024

DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image Personalization

CVPR 2024poster

The objective of text-to-image (T2I) personalization is to customize a diffusion model to a user-provided reference concept generating diverse images of the concept aligned with the target prompts. Conventional methods representing the reference concepts using unique text embeddings often fail to ac…

2024

Local All-Pair Correspondence for Point Tracking

ECCV 2024poster

"We introduce , a highly accurate and efficient model designed for the task of tracking any point (TAP) across video sequences. Previous approaches in this task often rely on local 2D correlation maps to establish correspondences from a point in the query image to a local region in the target image,…

2022

Cost Aggregation with 4D Convolutional Swin Transformer for Few-Shot Segmentation

ECCV 2022poster

"We present a novel cost aggregation network, called Volumetric Aggregation with Transformers (VAT), for few-shot segmentation. The use of transformers can benefit correlation map aggregation through self-attention over a global receptive field. However, the tokenization of a correlation map for tra…