← Search

Xiaokun Feng

16 accepted papers

2026

Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos

ICML 2026poster

Without incurring significant computational overhead, train-free long video generation aims to enable foundation video generation models to produce longer videos. Frame-level autoregressive frameworks, e.g., FIFO-diffusion, offer the advantage of generating infinitely long videos with constant memor…

Cited by 0SourceScholar
2026

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation

CVPR 2026

Few-step generation has been a long-standing goal, with recent one-step generation methods exemplified by MeanFlow achieving remarkable results. Existing research on MeanFlow primarily focuses on class-to-image generation. However, an intuitive yet unexplored direction is to extend the condition fro

Cited by 0SourcecodeScholar
2026

ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints

AAAI 2026technical

Video generation models have achieved remarkable progress, particularly excelling in realistic scenarios; however, their performance degrades notably in imaginative scenarios. These prompts often involve rarely co-occurring concepts with long-distance semantic relationships, falling outside training

Cited by 0SourcePDFScholar
2026

LATENT TEMPORAL DISCREPANCY AS MOTION PRIOR: A LOSS-WEIGHTING STRATEGY FOR DYNAMIC FIDELITY IN T2V

ICASSP 2026oral

Video generation models have achieved notable progress in static scenarios, yet their performance in motion video generation remains limited, with quality degrading under drastic dynamic changes. This is due to noise disrupting temporal coherence and increasing the difficulty of learning dynamic reg…

Cited by 0SourcePDFScholar
2026

NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation

ICLR 2026poster

With the rapid development of foundation video generation technologies, long video generation models have exhibited promising research potential thanks to expanded content creation space. Recent studies reveal that the goal of long video generation tasks is not only to extend video duration but also…

Cited by 0SourceScholar
2026

Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation

AAAI 2026technical

Visual effects (VFX) are essential visual enhancements fundamental to modern cinematic production. Although video generation models offer cost-efficient solutions for VFX production, current methods are constrained by per-effect LoRA training, which limits generation to single effects. This fundamen

Cited by 0SourcePDFScholar
2026

S$^2$-Guidance: Stochastic Self-Guidance for Training-Free Enhancement of Diffusion Models

ICLR 2026poster

Classifier-free Guidance (CFG) is a widely used technique in modern diffusion models for generating high-quality samples. However, through an empirical analysis on both Gaussian mixture models with closed-form solutions and real-world data distributions, we observe a discrepancy between the suboptim…

Cited by 0SourcecodeScholar
2025

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking

ICCV 2025poster

Vision-language tracking aims to locate the target object in the video sequence using a template patch and a language description provided in the initial frame. To achieve robust tracking, especially in complex long-term scenarios that reflect real-world conditions as recently highlighted by MGIT, i…

2025

CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features

ICML 2025poster

Effectively modeling and utilizing spatiotemporal features from RGB and other modalities (e.g., depth, thermal, and event data, denoted as X) is the core of RGB-X tracker design. Existing methods often employ two parallel branches to separately process the RGB and X input streams, requiring the mod…

2025

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues

ICASSP 2025accepted

Vision-Language Tracking (VLT) aims to localize a target in video sequences using a visual template and language description. While textual cues enhance tracking potential, current datasets typically contain much more image data than text, limiting the ability of VLT methods to align the two modalit…

Cited by 0SourceScholar
2025

Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards

NeurIPS 2025poster

Chain of thought reasoning has demonstrated remarkable success in large language models, yet its adaptation to vision-language reasoning remains an open challenge with unclear best practices. Existing attempts typically employ reasoning chains at a coarse-grained level, which struggles to perform fi…

Cited by 0SourcecodeScholar
2025

VMBench: A Benchmark for Perception-Aligned Video Motion Generation

ICCV 2025poster

Video generation has advanced rapidly, improving evaluation methods, yet assessing video's motion remains a major challenge. Specifically, there are two key issues: 1) current motion metrics do not fully align with human perceptions; 2) the existing motion prompts are limited. Based these findings,…

2024

Beyond Accuracy: Tracking more like Human via Visual Search

NeurIPS 2024poster

Human visual search ability enables efficient and accurate tracking of an arbitrary moving target, which is a significant research interest in cognitive neuroscience. The recently proposed Central-Peripheral Dichotomy (CPD) theory sheds light on how humans effectively process visual information and…

2024

MemVLT: Vision-Language Tracking with Adaptive Memory-based Prompts

NeurIPS 2024poster

Vision-language tracking (VLT) enhances traditional visual object tracking by integrating language descriptions, requiring the tracker to flexibly understand complex and diverse text in addition to visual information. However, most existing vision-language trackers still overly rely on initial fixed…

Cited by 5SourcePDFScholar
2024

Revealing the Dark Secrets of Extremely Large Kernel ConvNets on Robustness

ICML 2024poster

Robustness is a vital aspect to consider when deploying deep learning models into the wild. Numerous studies have been dedicated to the study of the robustness of vision transformers (ViTs), which have dominated as the mainstream backbone choice for vision tasks since the dawn of 2020s. Recently, so…

2023

A Multi-modal Global Instance Tracking Benchmark (MGIT): Better Locating Target in Complex Spatio-temporal and Causal Relationship

NeurIPS 2023poster

Tracking an arbitrary moving target in a video sequence is the foundation for high-level tasks like video understanding. Although existing visual-based trackers have demonstrated good tracking capabilities in short video sequences, they always perform poorly in complex environments, as represented b…

Cited by 13SourcePDFScholar