← Search

Xiaoming Wei

32 accepted papers

2026

Active Intelligence in Video Avatars via Closed-loop World Modeling

CVPR 2026

Current video avatar generation methods excel at identity preservation and motion alignment but lack genuine agency--they cannot autonomously pursue long-term goals through adaptive environmental interaction. We address this by introducing L-IVA (Long-horizon Interactive Visual Avatar), a task and b

Cited by 0SourceScholar
2026

Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory

ICML 2026poster

We propose **Infinite-World**, a robust interactive world model capable of maintaining coherent visual memory over **1000+ frames** in complex real-world environments. While existing world models can be efficiently optimized on synthetic data with perfect ground-truth, they lack an effective trainin…

Cited by 0SourceScholar
2026

LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts

ICML 2026poster

Recent advances in video diffusion models have significantly improved visual quality, yet ultra-high-resolution (UHR) video generation remains a formidable challenge due to the compounded difficulties of motion modeling, semantic planning, and detail synthesis. To address these limitations, we propo…

Cited by 0SourceScholar
2026

PositionIC: Unified Position and Identity Consistency for Image Customization

CVPR 2026

Recent subject-driven image customization excels in fidelity, yet fine-grained instance-level spatial control remains an elusive challenge, hindering real-world applications. This limitation stems from two factors: a scarcity of scalable, position-annotated datasets, and the entanglement of identity

Cited by 0SourcecodeScholar
2026

PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework

ICLR 2026poster

Generating aesthetic posters is more challenging than simple design images: it requires not only precise text rendering but also the seamless integration of abstract artistic content, striking layouts, and overall stylistic harmony. To address this, we propose PosterCraft, a unified framework that a…

Cited by 0SourcecodeScholar
2026

PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback

CVPR 2026

Image-to-poster generation is a high-demand task requiring not only local adjustments but also high-level design understanding. Models must generate text, layout, style, and visual elements while preserving semantic fidelity and aesthetic coherence. The process spans two regimes: local editing, wher

Cited by 0SourcecodeScholar
2026

PosterReward: Unlocking Accurate Evaluation for High-Quality Graphic Design Generation

CVPR 2026

Recent advancements in the text-rendering capabilities of image generation models have made the end-to-end creation of graphic design content, such as posters, increasingly feasible. However, existing reward models fall short of accurately assessing design quality, as they primarily focus on global

Cited by 0SourcecodeScholar
2026

U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation

CVPR 2026

Full-stack multimodal interaction in real-time is a central goal in building intelligent embodied agents capable of natural, dynamic communication. However, existing systems are either limited to unimodal generation or suffer from degraded reasoning and poor cross-modal alignment, preventing coheren

Cited by 0SourcecodeScholar
2026

UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models

CVPR 2026

With the advancement of multi-modal Large Language Models (LLMs), Video LLMs have been further developed to perform on holistic and specialized video understanding. However, existing works are limited to specialized video understanding tasks, failing to achieve a comprehensive and multi-grained vide

Cited by 0SourceScholar
2026

ViType: High-Fidelity Visual Text Rendering via Glyph-Aware Multimodal Diffusion

AAAI 2026technical

Current text-to-image models face challenges in visual text rendering: text encoders like CLIP and T5 lack glyph-level understanding and often struggle to distinguish between the specific words to be rendered and their intended semantic meaning within prompts. In addition, inconsistencies between th

Cited by 0SourcePDFScholar
2026

WildActor: Unconstrained Identity-Preserving Video Generation

ICML 2026poster

Production-ready human video generation requires digital actors to maintain strictly consistent full-body identities across dynamic shots, viewpoints and motions, a setting that remains challenging for existing methods. Prior methods often suffer from face-centric behavior that neglects body-level c…

Cited by 0SourceScholar
2025

ARIG: Autoregressive Interactive Head Generation for Real-time Conversations

ICCV 2025poster

Face-to-face communication, as a common human activity, motivates the research on interactive head generation. A virtual agent can generate motion responses with both listening and speaking capabilities based on the audio or motion signals of the other user and itself. However, previous clip-wise ge…

2025

LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

CVPR 2025poster

Recent advancements in multimodal large language models (MLLMs) have shown promising results, yet existing approaches struggle to effectively handle both temporal and spatial localization simultaneously. This challenge stems from two key issues: first, incorporating spatial-temporal localization int…

2025

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

NeurIPS 2025poster

Audio-driven human animation methods, such as talking head and talking body generation, have made remarkable progress in generating synchronized facial movements and appealing visual quality videos. However, existing methods primarily focus on single human animation and struggle with multi-stream au…

Cited by 0SourcecodeScholar
2025

Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation

AAAI 2025technical

In this paper, we propose an Audio-Language-Referenced SAM 2 (AL-Ref-SAM 2) pipeline to explore the training-free paradigm for audio and language-referenced video object segmentation, namely AVS and RVOS tasks. The intuitive solution leverages GroundingDINO to identify the target object from a singl…

2024

BEM: Balanced and Entropy-based Mix for Long-Tailed Semi-Supervised Learning

CVPR 2024poster

Data mixing methods play a crucial role in semi-supervised learning (SSL) but their application is unexplored in long-tailed semi-supervised learning (LTSSL). The primary reason is that the in-batch mixing manner fails to address class imbalance. Furthermore existing LTSSL methods mainly focus on re…

Cited by 7SourcePDFScholar
2024

ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and Spotting

CVPR 2024poster

Abstract In recent years text-image joint pre-training techniques have shown promising results in various tasks. However in Optical Character Recognition (OCR) tasks aligning text instances with their corresponding text regions in images poses a challenge as it requires effective alignment between t…

2023

Bridging Search Region Interaction With Template for RGB-T Tracking

CVPR 2023poster

RGB-T tracking aims to leverage the mutual enhancement and complement ability of RGB and TIR modalities for improving the tracking process in various scenarios, where cross-modal interaction is the key component. Some previous methods concatenate the RGB and TIR search region features directly to pe…

2023

Elastic Aggregation for Federated Optimization

CVPR 2023poster

Federated learning enables the privacy-preserving training of neural network models using real-world data across distributed clients. FedAvg has become the preferred optimizer for federated learning because of its simplicity and effectiveness. FedAvg uses naive aggregation to update the server model…

2023

Enriching Phrases with Coupled Pixel and Object Contexts for Panoptic Narrative Grounding

IJCAI 2023poster

Panoptic narrative grounding (PNG) aims to segment things and stuff objects in an image described by noun phrases of a narrative caption. As a multimodal task, an essential aspect of PNG is the visual-linguistic interaction between image and caption. The previous two-stage method aggregates visual c…

Cited by 5SourcePDFScholar
2023

Masked Auto-Encoders Meet Generative Adversarial Networks and Beyond

CVPR 2023poster

Masked Auto-Encoder (MAE) pretraining methods randomly mask image patches and then train a vision Transformer to reconstruct the original pixels based on the unmasked patches. While they demonstrates impressive performance for downstream vision tasks, it generally requires a large amount of training…

Cited by 20SourcePDFScholar
2023

Rethinking skip connection model as a learnable Markov chain

ICLR 2023poster

Over the past few years afterward the birth of ResNet, skip connection has become the defacto standard for the design of modern architectures due to its widespread adoption, easy optimization, and proven performance. Prior work has explained the effectiveness of the skip connection mechanism from di…

2023

Uncertainty-Aware Image Captioning

AAAI 2023technical

It is well believed that the higher uncertainty in a word of the caption, the more inter-correlated context information is required to determine it. However, current image captioning methods usually consider the generation of all words in a sentence sequentially and equally. In this paper, we propos…

Cited by 19SourcePDFScholar
2022

Adaptive Spatial-BCE Loss for Weakly Supervised Semantic Segmentation

ECCV 2022poster

"For Weakly-Supervised Semantic Segmentation (WSSS) with image-level annotation, mostly relies on the classification network to generate initial segmentation pseudo-labels. However, the optimization target of classification networks usually neglects the discrimination between different pixels, like…

2022

Language-Bridged Spatial-Temporal Interaction for Referring Video Object Segmentation

CVPR 2022poster

Referring video object segmentation aims to predict foreground labels for objects referred by natural language expressions in videos. Previous methods either depend on 3D ConvNets or incorporate additional 2D ConvNets as encoders to extract mixed spatial-temporal features. However, these methods suf…

Cited by 74PDFcodeScholar
2022

Rethinking the Optimization of Average Precision: Only Penalizing Negative Instances before Positive Ones Is Enough

AAAI 2022technical

Optimising the approximation of Average Precision (AP) has been widely studied for image retrieval. Limited by the definition of AP, such methods consider both negative and positive instances ranking before each positive instance. However, we claim that only penalizing negative instances before posi…

2021

Embedded Discriminative Attention Mechanism for Weakly Supervised Semantic Segmentation

CVPR 2021poster

Weakly Supervised Semantic Segmentation (WSSS) with image-level annotation uses class activation maps from the classifier as pseudo-labels for semantic segmentation. However, such activation maps usually highlight the local discriminative regions rather than the whole object, which deviates from the…

Cited by 178PDFcodeScholar
2021

Rethinking BiSeNet for Real-Time Semantic Segmentation

CVPR 2021poster

BiSeNet has been proved to be a popular two-stream network for real-time segmentation. However, its principle of adding an extra path to encode spatial information is time-consuming, and the backbones borrowed from pretrained tasks, e.g., image classification, may be inefficient for image segmentati…

Cited by 814PDFcodeScholar