← Search

Xiaodong Wang

26 accepted papers

2026

A Training-Free Framework for Long Video Understanding via Video-Query-Options Similarity

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved remarkable success in image and short video understanding tasks, but their performance on hour-long videos remains limited due to constraint of input token capacity. Existing approaches often require costly training procedures, hindering their a…

Cited by 0SourceScholar
2026

Fine-flow Distilling Coarse-flow Video Generation for Long-Term Driving World Model

AAAI 2026technical

Driving world models are used to simulate futures by video generation based on the condition of the current state and actions. However, current models often suffer serious error accumulations when predicting the long-term future, which limits practical applications. Recent studies utilize the Diffu

Cited by 0SourcePDFScholar
2026

Incentivizing Versatile Video Reasoning in MLLMs via Data-Efficient Reinforcement Learning

CVPR 2026

Multimodal Large Language Models (MLLMs) have made great progress in video understanding tasks. However, when it comes to understanding complex or lengthy videos, MLLMs tend to overlook details or produce hallucinations. To alleviate these issues, recent work has attempted to leverage reinforcement

Cited by 0SourcecodeScholar
2026

Joint Spectral Image Reconstruction and Semantic Segmentation with Cooperative Unfolding

CVPR 2026

Coded Aperture Snapshot Spectral Imaging (CASSI) is an emerging hyperspectral image (HSI) acquisition technique for downstream semantic segmentation. Due to the ill-posedness nature of CASSI systems, typical solutions are compelled to conduct a two-stage reconstruction-then-segmentation pipeline, na

Cited by 0SourcecodeScholar
2026

LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding

AAAI 2026technical

The development of multimodal large language models (MLLMs) has advanced general video understanding. However, existing video evaluation benchmarks primarily focus on non-interactive videos, such as movies and recordings. To fill this gap, this paper proposes the first omnimodal benchmark for intera

Cited by 0SourcePDFScholar
2026

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models

CVPR 2026

Multi-modal large language models (MLLMs) have rapidly advanced in visual tasks, yet their spatial understanding remains limited to single images, leaving them ill-suited for physical-world applications that require multi-frame reasoning. In this paper, we propose a framework to equip MLLMs with mul

Cited by 0SourcecodeScholar
2026

TI-3DGS: 3D Thermal Reconstruction Via Thermal Imaging-Guided 3D Gaussian Splatting

ICRA 2026poster

Thermal imaging, with its all-weather capabilities and strong penetration, enables 3D reconstruction in low- light and adverse conditions. In this paper, we investigate RGB-independent pure 3D thermal reconstruction, aiming to overcome the challenges of 3D reconstruction in extreme environments wher…

Cited by 0Scholar
2025

EFTViT: Efficient Federated Training of Vision Transformers with Masked Images on Resource-Constrained Clients

ICCV 2025poster

Federated learning research has recently shifted from Convolutional Neural Networks (CNNs) to Vision Transformers (ViTs) due to their superior capacity. ViTs training demands higher computational resources due to the lack of 2D inductive biases inherent in CNNs. However, efficient federated training…

Cited by 0SourcePDFScholar
2025

EffiQA: Efficient Question-Answering with Strategic Multi-Model Collaboration on Knowledge Graphs

COLING 2025main

While large language models (LLMs) have shown remarkable capabilities in natural language processing, they struggle with complex, multi-step reasoning tasks involving knowledge graphs (KGs). Existing approaches that integrate LLMs and KGs either underutilize the reasoning abilities of LLMs or suffer…

Cited by 4SourcePDFScholar
2025

FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation

ICCV 2025poster

Vision-Language-Action (VLA) models have significantly advanced robotic manipulation by enabling robots to interpret language instructions for task execution. However, training these models often relies on large-scale user-specific data, raising concerns about privacy and security, which in turn lim…

Cited by 0SourcePDFScholar
2025

Proximal Algorithm Unrolling: Flexible and Efficient Reconstruction Networks for Single-Pixel Imaging

CVPR 2025poster

Deep-unrolling and plug-and-play (PnP) approaches have become the de-facto standard solvers for single-pixel imaging (SPI) inverse problem. PnP approaches, a class of iterative algorithms where regularization is implicitly performed by an off-the-shelf deep denoiser, are flexible for varying compres…

2025

SCI-Gaussian: Optimizing 3D Gaussian Radiance Fields from a Snapshot Compressive Image

ICASSP 2025accepted

Snapshot compressive imaging (SCI) is a compressed sensing (CS)-based high-speed imaging modality. Recent efforts have explored the underlying 3D representation from only an SCI image using neural radiance fields (NeRF), yet the training time, rendering computation cost, and reconstruction quality l…

Cited by 0SourceScholar
2025

Spectral Compressive Imaging via Chromaticity-Intensity Decomposition

NeurIPS 2025poster

In coded aperture snapshot spectral imaging (CASSI), the captured measurement entangles spatial and spectral information, posing a severely ill-posed inverse problem for hyperspectral images (HSIs) reconstruction. Moreover, the captured radiance inherently depends on scene illumination, making it di…

Cited by 0SourcecodeScholar
2024

Learning Invariant Representation with Consistency and Diversity for Semi-Supervised Source Hypothesis Transfer

ICASSP 2024accepted

Semi-supervised Domain adaptation (SSDA) has shown promising results by leveraging unlabeled data and limited labeled samples in the target domain. However, accessibility to source data is hindered by data privacy concerns, giving rise to Semi-supervised Source Hypothesis Transfer (SSHT). Integratin…

Cited by 0SourceScholar
2024

ORES: Open-Vocabulary Responsible Visual Synthesis

AAAI 2024technical

Avoiding synthesizing specific visual concepts is an essential challenge in responsible visual synthesis. However, the visual concept that needs to be avoided for responsible visual synthesis tends to be diverse, depending on the region, context, and usage scenarios. In this work, we formalize a new…

2024

SCINeRF: Neural Radiance Fields from a Snapshot Compressive Image

CVPR 2024highlight

In this paper we explore the potential of Snapshot Com- pressive Imaging (SCI) technique for recovering the under- lying 3D scene representation from a single temporal com- pressed image. SCI is a cost-effective method that enables the recording of high-dimensional data such as hyperspec- tral or te…

2023

Learning 3D Photography Videos via Self-supervised Diffusion on Single Images

IJCAI 2023poster

3D photography renders a static image into a video with appealing 3D visual effects. Existing approaches typically first conduct monocular depth estimation, then render the input frame to subsequent frames with various viewpoints, and finally use an inpainting model to fill those missing/occluded re…

Cited by 4SourcePDFScholar
2023

NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation

ACL 2023long

In this paper, we propose NUWA-XL, a novel Diffusion over Diffusion architecture for eXtremely Long video generation. Most current work generates long videos segment by segment sequentially, which normally leads to the gap between training on short videos and inferring long videos, and the sequentia…

Cited by 118SourcePDFScholar
2023

Progressive Meta-Pooling Learning for Lightweight Image Classification Model

ICASSP 2023accepted

Practical networks for edge devices adopt shallow depth and small convolutional kernels to save memory and computational cost, which leads to a restricted receptive field. Conventional efficient learning methods focus on lightweight convolution designs, ignoring the role of the receptive field in ne…

Cited by 0SourceScholar
2023

RD-NAS: Enhancing One-Shot Supernet Ranking Ability Via Ranking Distillation From Zero-Cost Proxies

ICASSP 2023accepted

Neural architecture search (NAS) has made tremendous progress in the automatic design of effective neural network structures but suffers from a heavy computational burden. One-shot NAS significantly alleviates the burden through weight sharing and improves computational efficiency. Zero-shot NAS fur…

Cited by 0SourceScholar
2022

RSGT: Relational Structure Guided Temporal Relation Extraction

COLING 2022main

Temporal relation extraction aims to extract temporal relations between event pairs, which is crucial for natural language understanding. Few efforts have been devoted to capturing the global features. In this paper, we propose RSGT: Relational Structure Guided Temporal Relation Extraction to extrac…

Cited by 25SourcePDFScholar
2021

Towards Extremely Compact RNNs for Video Recognition With Fully Decomposed Hierarchical Tucker Structure

CVPR 2021poster

Recurrent Neural Networks (RNNs) have been widely used in sequence analysis and modeling. However, when processing high-dimensional data, RNNs typically require very large model sizes, thereby bringing a series of deployment challenges. Although various prior works have been proposed to reduce the R…

Cited by 38PDFScholar
2016

Tensor completion via adaptive sampling of tensor fibers: Application to efficient indoor RF fingerprinting

ICASSP 2016accepted

In this paper, we consider tensor completion under adaptive sampling of tensor (a multidimensional array) fibers. Tensor fibers or tubes are vectors obtained by fixing all but one index of the array. This sampling is in contrast to the cases considered so far where one performs an adaptive element-w…

Cited by 0SourceScholar