← Search

Meiqi Wu

15 accepted papers

2026

ARTIFACT-AWARE EVALUATION FOR HIGH-QUALITY VIDEO GENERATION

ICASSP 2026oral

With the rapid advancement of video generation techniques, evaluating and auditing generated videos has become increasingly crucial. Existing approaches typically offer coarse video quality scores, lacking detailed localization and categorization of specific artifacts. In this work, we introduce a c…

Cited by 0SourcePDFScholar
2026

Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos

ICML 2026poster

Without incurring significant computational overhead, train-free long video generation aims to enable foundation video generation models to produce longer videos. Frame-level autoregressive frameworks, e.g., FIFO-diffusion, offer the advantage of generating infinitely long videos with constant memor…

Cited by 0SourceScholar
2026

ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints

AAAI 2026technical

Video generation models have achieved remarkable progress, particularly excelling in realistic scenarios; however, their performance degrades notably in imaginative scenarios. These prompts often involve rarely co-occurring concepts with long-distance semantic relationships, falling outside training

Cited by 0SourcePDFScholar
2026

LATENT TEMPORAL DISCREPANCY AS MOTION PRIOR: A LOSS-WEIGHTING STRATEGY FOR DYNAMIC FIDELITY IN T2V

ICASSP 2026oral

Video generation models have achieved notable progress in static scenarios, yet their performance in motion video generation remains limited, with quality degrading under drastic dynamic changes. This is due to noise disrupting temporal coherence and increasing the difficulty of learning dynamic reg…

Cited by 0SourcePDFScholar
2026

NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation

ICLR 2026poster

With the rapid development of foundation video generation technologies, long video generation models have exhibited promising research potential thanks to expanded content creation space. Recent studies reveal that the goal of long video generation tasks is not only to extend video duration but also…

Cited by 0SourceScholar
2026

Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation

AAAI 2026technical

Visual effects (VFX) are essential visual enhancements fundamental to modern cinematic production. Although video generation models offer cost-efficient solutions for VFX production, current methods are constrained by per-effect LoRA training, which limits generation to single effects. This fundamen

Cited by 0SourcePDFScholar
2026

S$^2$-Guidance: Stochastic Self-Guidance for Training-Free Enhancement of Diffusion Models

ICLR 2026poster

Classifier-free Guidance (CFG) is a widely used technique in modern diffusion models for generating high-quality samples. However, through an empirical analysis on both Gaussian mixture models with closed-form solutions and real-world data distributions, we observe a discrepancy between the suboptim…

Cited by 0SourcecodeScholar
2026

Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement Learning

CVPR 2026

Recent studies have demonstrated significant progress in aligning text-to-image diffusion models with human preference via Reinforcement Learning from Human Feedback. However, while existing methods achieve high scores on automated reward metrics, they often lead to Preference Mode Collapse (PMC)-a

Cited by 0SourceScholar
2025

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking

ICCV 2025poster

Vision-language tracking aims to locate the target object in the video sequence using a template patch and a language description provided in the initial frame. To achieve robust tracking, especially in complex long-term scenarios that reflect real-world conditions as recently highlighted by MGIT, i…

2025

CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features

ICML 2025poster

Effectively modeling and utilizing spatiotemporal features from RGB and other modalities (e.g., depth, thermal, and event data, denoted as X) is the core of RGB-X tracker design. Existing methods often employ two parallel branches to separately process the RGB and X input streams, requiring the mod…

2025

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues

ICASSP 2025accepted

Vision-Language Tracking (VLT) aims to localize a target in video sequences using a visual template and language description. While textual cues enhance tracking potential, current datasets typically contain much more image data than text, limiting the ability of VLT methods to align the two modalit…

Cited by 0SourceScholar
2025

VMBench: A Benchmark for Perception-Aligned Video Motion Generation

ICCV 2025poster

Video generation has advanced rapidly, improving evaluation methods, yet assessing video's motion remains a major challenge. Specifically, there are two key issues: 1) current motion metrics do not fully align with human perceptions; 2) the existing motion prompts are limited. Based these findings,…

2024

Beyond Accuracy: Tracking more like Human via Visual Search

NeurIPS 2024poster

Human visual search ability enables efficient and accurate tracking of an arbitrary moving target, which is a significant research interest in cognitive neuroscience. The recently proposed Central-Peripheral Dichotomy (CPD) theory sheds light on how humans effectively process visual information and…

2024

MemVLT: Vision-Language Tracking with Adaptive Memory-based Prompts

NeurIPS 2024poster

Vision-language tracking (VLT) enhances traditional visual object tracking by integrating language descriptions, requiring the tracker to flexibly understand complex and diverse text in addition to visual information. However, most existing vision-language trackers still overly rely on initial fixed…

Cited by 5SourcePDFScholar
2023

A Multi-modal Global Instance Tracking Benchmark (MGIT): Better Locating Target in Complex Spatio-temporal and Causal Relationship

NeurIPS 2023poster

Tracking an arbitrary moving target in a video sequence is the foundation for high-level tasks like video understanding. Although existing visual-based trackers have demonstrated good tracking capabilities in short video sequences, they always perform poorly in complex environments, as represented b…

Cited by 13SourcePDFScholar