← Search

Xun Guo

14 accepted papers

2026

MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement

ICLR 2026poster

We tackle the task of any-reference video generation, which aims to synthesize videos conditioned on arbitrary types and combinations of reference subjects, together with textual prompts. This task faces persistent challenges, including identity inconsistency, entanglement among multiple reference s…

Cited by 0SourcecodeScholar
2025

ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation

ICLR 2025poster

Text-to-video (T2V) models have recently undergone rapid and substantial advancements. Nevertheless, due to limitations in data and computational resources, achieving efficient generation of long videos with rich motion dynamics remains a significant challenge. To generate high-quality, dynamic, an…

Cited by 6SourcePDFScholar
2025

I2VGuard: Safeguarding Images against Misuse in Diffusion-based Image-to-Video Models

CVPR 2025poster

Recent advances in image-to-video generation have enabled animation of still images and offered pixel-level controllability. While these models hold great potential to transform single images into vivid and dynamic videos, they also carry risks of misuse that could impact privacy, security, and copy…

Cited by 0SourcePDFScholar
2025

Image as a World: Generating Interactive World from Single Image via Panoramic Video Generation

NeurIPS 2025poster

Generating an interactive visual world from a single image is both challenging and practically valuable, as single-view inputs are easy to acquire and align well with prompt-driven applications such as gaming and virtual reality. This paper introduces a novel unified framework, Image as a World (**I…

Cited by 0SourceScholar
2025

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

CVPR 2025poster

This work explores whether a deep generative model can learn complex knowledge solely from visual input, in contrast to the prevalent focus on text-based models like large language models (LLMs). We develop VideoWorld, an auto-regressive video generation model trained on unlabeled video data, and te…

Cited by 8SourcePDFScholar
2024

DeTeCtive: Detecting AI-generated Text via Multi-Level Contrastive Learning

NeurIPS 2024poster

Current techniques for detecting AI-generated text are largely confined to manual feature crafting and supervised binary classification paradigms. These methodologies typically lead to performance bottlenecks and unsatisfactory generalizability. Consequently, these methods are often inapplicable for…

2024

MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

CVPR 2024poster

Recently integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet existing systems can only handle videos with very few frames. For long videos the computation complexity memory cost and…

2024

Noise-assisted Prompt Learning for Image Forgery Detection and Localization

ECCV 2024poster

"We present CLIP-IFDL, a novel image forgery detection and localization (IFDL) model that harnesses the power of Contrastive Language Image Pre-Training (CLIP). However, directly incorporating CLIP in forgery detection poses challenges, given its lack of specific prompts and forgery consciousness. T…

Cited by 4SourcePDFScholar