← Search

Yuanfan Guo

11 accepted papers

2025

Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising

ICASSP 2025accepted

Recent advances in diffusion models have greatly improved text-driven video generation. However, training models for long video generation demands significant computational power and extensive data, leading most video diffusion models to be limited to a small number of frames. Existing training-free…

Cited by 0SourceScholar
2025

DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance

ICASSP 2025accepted

Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods frequently integrate semantic information from images or simply concatenate images, which often leads to low fidelity and flickering in the generated videos. T…

Cited by 25SourceScholar
2025

EasyControl: Adding Control to Video Diffusion for Controllable Video Generation and Interpolation

ICASSP 2025accepted

The diffusion model is widely leveraged for either controllable video generation or video interpolation. As each field has its task-specific problems, it is difficult to merely develop a single model for completing both tasks simultaneously. Moreover, most existing works only support image condition…

Cited by 0SourceScholar
2025

FreqPrior: Improving Video Diffusion Models with Frequency Filtering Gaussian Noise

ICLR 2025poster

Text-driven video generation has advanced significantly due to developments in diffusion models. Beyond the training and sampling phases, recent studies have investigated noise priors of diffusion models, as improved noise priors yield better generation results. One recent approach employs the Fouri…

Cited by 0SourcePDFScholar
2025

VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning

ICCV 2025poster

In recent years, video question answering based on multimodal large language models (MLLM) has garnered considerable attention, due to the benefits from the substantial advancements in LLMs. However, these models have a notable deficiency in the domains of video temporal grounding and reasoning, pos…

2024

Any-Size-Diffusion: Toward Efficient Text-Driven Synthesis for Any-Size HD Images

AAAI 2024technical

Stable diffusion, a generative model used in text-to-image synthesis, frequently encounters resolution-induced composition problems when generating images of varying sizes. This issue primarily stems from the model being trained on pairs of single-scale images and their corresponding text descriptio…

2024

HumanRefiner: Benchmarking Abnormal Human Generation and Refining with Coarse-to-fine Pose-Reversible Guidance

ECCV 2024poster

"Text-to-image diffusion models have significantly advanced in conditional image generation. However, these models usually struggle with accurately rendering images featuring humans, resulting in distorted limbs and other anomalies. This issue primarily stems from the insufficient recognition and ev…

2024

PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion

ECCV 2024poster

"Current large-scale diffusion models represent a giant leap forward in conditional image synthesis, capable of interpreting diverse cues like text, human poses, and edges. However, their reliance on substantial computational resources and extensive data collection remains a bottleneck. On the other…

2024

Self-Adaptive Reality-Guided Diffusion for Artifact-Free Super-Resolution

CVPR 2024poster

Artifact-free super-resolution (SR) aims to translate low-resolution images into their high-resolution counterparts with a strict integrity of the original content eliminating any distortions or synthetic details. While traditional diffusion-based SR techniques have demonstrated remarkable abilities…

2024

SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM

NeurIPS 2024poster

Large language models (LLMs) have demonstrated exceptional capabilities in text understanding, which has paved the way for their expansion into video LLMs (Vid-LLMs) to analyze video data. However, current Vid-LLMs struggle to simultaneously retain high-quality frame-level semantic information (i.e.…

Cited by 2SourcePDFScholar
2022

HCSC: Hierarchical Contrastive Selective Coding

CVPR 2022poster

Hierarchical semantic structures naturally exist in an image dataset, in which several semantically relevant image clusters can be further integrated into a larger cluster with coarser-grained semantics. Capturing such structures with image representations can greatly benefit the semantic understand…

Cited by 101PDFcodeScholar