← Search

Weining Wang

11 accepted papers

2026

Divid: Disentangled Spatial-Temporal Modeling within LLMs for Temporally Grounded Video Understanding

ICLR 2026poster

Recent advances in Video LLMs have improved video understanding performance, but temporally grounded understanding in long-form videos remains challenging. Most models encode video frames into a flat sequence of visual tokens, which are then processed together with textual input by the LLM. While ef…

Cited by 0SourceScholar
2026

TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Models

CVPR 2026

Vision-Language Models (VLMs), such as CLIP, have achieved impressive zero-shot recognition performance but remain highly susceptible to adversarial perturbations, posing significant risks in safety-critical scenarios. Previous training-time defenses rely on adversarial fine-tuning, which requires l

Cited by 0SourcecodeScholar
2026

UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception

AAAI 2026technical

The remarkable success of diffusion models in text-to-image generation has sparked growing interest in expanding their capabilities to a variety of multi-modal tasks, including image understanding, manipulation, and perception. These tasks require advanced semantic comprehension across both visual a

Cited by 0SourcePDFScholar
2026

VisualPrompter: Semantic-Aware Prompt Optimization with Visual Feedback for Text-to-Image Synthesis

ICLR 2026poster

The notable gap between user-provided and model-preferred prompts poses a significant challenge for generating high-quality images with text-to-image models, compelling the need for prompt engineering. Current studies on prompt engineering can effectively enhance the style and aesthetics of generate…

Cited by 0SourcecodeScholar
2026

W-EDIT: A Wavelet-Based Frequency-Aware Framework for Text-Driven Image Editing

ICLR 2026poster

While recent advances in Diffusion Transformers (DiTs) have significantly advanced text-to-image generation, text-driven image editing remains challenging. Existing approaches either struggle to balance structural preservation with flexible modifications or require costly fine-tuning of large models…

Cited by 0SourceScholar
2025

AR-Diffusion: Asynchronous Video Generation with Auto-Regressive Diffusion

CVPR 2025poster

The task of video generation requires synthesizing visually realistic and temporally coherent video frames. Existing methods primarily use asynchronous auto-regressive models or synchronous diffusion models to address this challenge. However, asynchronous auto-regressive models often suffer from inc…

2023

GLOBER: Coherent Non-autoregressive Video Generation via GLOBal Guided Video DecodER

NeurIPS 2023poster

Video generation necessitates both global coherence and local realism. This work presents a novel non-autoregressive method GLOBER, which first generates global features to obtain comprehensive global guidance and then synthesizes video frames based on the global features to generate coherent videos…

2023

MOSO: Decomposing MOtion, Scene and Object for Video Prediction

CVPR 2023poster

Motion, scene and object are three primary visual components of a video. In particular, objects represent the foreground, scenes represent the background, and motion traces their dynamics. Based on this insight, we propose a two-stage MOtion, Scene and Object decomposition framework (MOSO) for video…

2023

WL-MSR: Watch and Listen for Multimodal Subtitle Recognition

ICASSP 2023accepted

Video subtitles could be defined as the combination of visualized subtitles in frames and textual content recognized from speech, which play a significant role in video understanding for both humans and machines. In this paper, we propose a novel Watch and Listen for Multimodal Subtitle Recognition…

Cited by 0SourceScholar
2021

HAIR: Hierarchical Visual-Semantic Relational Reasoning for Video Question Answering

ICCV 2021poster

Relational reasoning is at the heart of video question answering. However, existing approaches suffer from several common limitations: (1) they only focus on either object-level or frame-level relational reasoning, and fail to integrate the both; and (2) they neglect to leverage semantic knowledge f…

Cited by 62PDFcodeScholar
2019

Language-Driven Temporal Activity Localization: A Semantic Matching Reinforcement Learning Model

CVPR 2019oral

Current studies on action detection in untrimmed videos are mostly designed for action classes, where an action is described at word level such as jumping, tumbling, swing, etc. This paper focuses on a rarely investigated problem of localizing an activity via a sentence query which would be more cha…

Cited by 215PDFScholar