← Search

Cunjian Chen

10 accepted papers

2026

DynaMem: Consistent Long Video Generation via Hierarchical Memory and Motion Priors

ICML 2026poster

Recent text-to-video diffusion models can synthesize visually compelling clips from natural language prompts. However, practical applications increasingly demand long-form videos with evolving narratives and persistent identity. A common solution is autoregressive generation, where the video is prod…

Cited by 0SourceScholar
2026

Fresco: Frequency-Spatial Consistent Optimization for Fine-Grained Head Avatar Modeling

CVPR 2026

We propose Fresco, a unified optimization pipeline designed to mitigate early over-sharpening, and cross-view drifting in head avatar reconstruction. Fresco combines a Laplacian-pyramid-based frequency curriculum with UV-space consistency regularization to progressively enhance reconstruction qualit

Cited by 0SourcecodeScholar
2026

IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation

AAAI 2026technical

Recent visual generative models enable story generation with consistent characters from text, but human-centric story generation faces additional challenges, such as maintaining detailed and diverse human face consistency and coordinating multiple characters across different images. This paper prese

Cited by 0SourcePDFScholar
2026

Less is More: Geometric Unlearning for LLMs with Minimal Data Disclosure

ICML 2026poster

As large language models (LLMs) are increasingly deployed in real-world systems, they must support post-hoc removal of specific content to meet privacy and governance requirements. This motivates selective unlearning, which suppresses information about a particular entity or topic while preserving t…

Cited by 0SourceScholar
2026

Mamba-Driven Multi-View Discriminative Clustering via Global-Local Cross-View Sequence Modeling

AAAI 2026technical

Multi-view clustering (MVC) has recently garnered increasing attention for its ability to partition unlabeled samples into distinct clusters by leveraging complementary and consistent information from different views. Existing MVC methods primarily combine deep neural networks with contrastive learn

Cited by 0SourcePDFScholar
2026

OmniShow: Orchestrating Multimodal Conditions for Human-Object Interaction Video Generation

ICML 2026poster

In this work, we study **Human-Object Interaction Video Generation (HOIVG)**, which aims to synthesize high-quality HOI videos via text, reference image, audio, and pose conditions. To address the challenges of harmonious multimodal injection and heterogeneous data utility, we present **OmniShow**, …

Cited by 0SourceScholar
2026

Synthetic Curriculum Reinforces Compositional Text-to-Image Generation

CVPR 2026

Text-to-Image (T2I) generation has long been an open problem, with compositional synthesis remaining particularly challenging. This task requires accurate rendering of complex scenes containing multiple objects that exhibit diverse attributes as well as intricate spatial and semantic relationships,

Cited by 0SourceScholar
2025

Consistent and Controllable Image Animation with Motion Diffusion Models

CVPR 2025poster

Diffusion models have achieved significant progress in the task of image animation due to their powerful generative capabilities. However, preserving appearance consistency to the static input image, and avoiding abrupt motion change in the generated animation, remains challenging. In this paper, we…

Cited by 0SourcePDFScholar
2025

SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency

NeurIPS 2025poster

Recent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus on character consistency, they often overlook the crucial role of scenes in storytelling, which restricts their creativit…

Cited by 0SourceScholar
2025

Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting

NeurIPS 2025poster

The increase in computing power and the necessity of AI-assisted decision-making boost the growing application of large language models (LLMs). Along with this, the potential retention of sensitive data of LLMs has spurred increasing research into machine unlearning. However, existing unlearning app…

Cited by 0SourceScholar