← Search

Xijun Wang

25 accepted papers

2026

Beyond Test-Time Training: Learning to Reason via Hardware-Efficient Optimal Control

ICML 2026poster

Associative memory has long underpinned the design of sequential models. Beyond recall, humans reason by *projecting future states and selecting goal-directed actions*, a capability that modern language models increasingly require but do not natively encode. While prior work uses reinforcement learn…

Cited by 0SourceScholar
2026

Identifying and Evaluating Inactive Heads in Pretrained LLMs

ICLR 2026poster

Attention is foundational to large language models (LLMs), enabling different heads to have diverse focus on relevant input tokens. However, learned behaviors like attention sinks, where the first token receives the most attention despite limited semantic importance, suggest some heads may be inacti…

Cited by 0SourceScholar
2026

NewtonGen: Physics-consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics

ICLR 2026poster

A primary bottleneck in large-scale text-to-video generation today is physical consistency and controllability. Despite recent advances, state-of-the-art models often produce unrealistic motions, such as objects falling upward, or abrupt changes in velocity and direction. Moreover, these models lack…

Cited by 0SourcecodeScholar
2026

SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation

CVPR 2026

Images and videos are discrete 2D projections of the 4D world (3D space + time). Most visual understanding, prediction, and generation operate directly on 2D observations, leading to suboptimal performance. We propose SeeU, a novel approach that learns the continuous 4D dynamics and generate the uns

Cited by 0SourcecodeScholar
2025

Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis

CVPR 2025highlight

Image generation today can produce somewhat realistic images from text prompts. However, if one asks the generator to synthesize a specific camera setting such as creating different fields of view using a 24mm lens versus a 70mm lens, the generator will not be able to interpret and generate scene-co…

2025

Learning Phase Distortion with Selective State Space Models for Video Turbulence Mitigation

CVPR 2025highlight

Atmospheric turbulence is a major source of image degradation in long-range imaging systems. Although numerous deep learning-based turbulence mitigation (TM) methods have been proposed, many are slow, memory-hungry, and do not generalize well. In the spatial domain, methods based on convolutional op…

2025

WorldSimBench: Towards Video Generation Models as World Simulators

ICML 2025poster

Recent advancements in predictive models have demonstrated exceptional capabilities in predicting the future state of objects and scenes. However, the lack of categorization based on inherent characteristics continues to hinder the progress of predictive model development. Additionally, existing ben…

Cited by 18SourcePDFScholar
2024

AGL-Net: Aerial-Ground Cross-Modal Global Localization with Varying Scales

IROS 2024poster

We present AGL-NET, a novel learning-based method for global localization using LiDAR point clouds and satellite maps. AGL-Net tackles two critical challenges: bridging the representation gap between image and points modalities for robust feature matching, and handling inherent scale discrepancies b…

Cited by 1SourcecodeScholar
2024

Adv-Diffusion: Imperceptible Adversarial Face Identity Attack via Latent Diffusion Model

AAAI 2024technical

Adversarial attacks involve adding perturbations to the source image to cause misclassification by the target model, which demonstrates the potential of attacking face recognition models. Existing adversarial face image generation methods still can’t achieve satisfactory performance because of low t…

2024

AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models

EMNLP 2024finding

Large vision-language models (LVLMs) are prone to hallucinations, where certain contextual cues in an image can trigger the language module to produce overconfident and incorrect reasoning about abnormal or hypothetical objects. While some benchmarks have been developed to investigate LVLM hallucina…

2024

Deep Stochastic Kinematic Models for Probabilistic Motion Forecasting in Traffic

IROS 2024poster

In trajectory forecasting tasks for traffic, future output trajectories can be computed by advancing the ego vehicle’s state with predicted actions according to a kinematics model. By unrolling predicted trajectories via time integration and models of kinematic dynamics, predicted trajectories shoul…

Cited by 0SourceScholar
2024

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

CVPR 2024poster

We introduce "HallusionBench" a comprehensive benchmark designed for the evaluation of image-context reasoning. This benchmark presents significant challenges to advanced large visual-language models (LVLMs) such as GPT-4V(ision) Gemini Pro Vision Claude 3 and LLaVA-1.5 by emphasizing nuanced unders…

2024

ICAR: Image-Based Complementary Auto Reasoning

AAAI 2024technical

Scene-aware Complementary Item Retrieval (CIR) is a challenging task which requires to generate a set of compatible items across domains. Due to the subjectivity, it is difficult to set up a rigorous standard for both data collection and learning objectives. To address this challenging task, we prop…

Cited by 1SourcePDFScholar
2024

SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition

IROS 2024poster

We present a new learning approach, Soft Conditional Prompt Learning (SCP), which leverages the strengths of prompt learning for aerial video action recognition. Our approach is designed to predict the action of each agent by helping the models focus on the descriptions or instructions associated wi…

Cited by 1SourceScholar
2023

AZTR: Aerial Video Action Recognition with Auto Zoom and Temporal Reasoning

ICRA 2023poster

We propose a novel approach for aerial video action recognition. Our method is designed for videos captured using UAVs and can run on edge or mobile devices. We present a learning-based approach that uses customized auto zoom to automatically identify the human target and scale it appropriately. Thi…

Cited by 18SourceScholar
2023

CrossLoc3D: Aerial-Ground Cross-Source 3D Place Recognition

ICCV 2023poster

We present CrossLoc3D, a novel 3D place recognition method that solves a large-scale point matching problem in a cross-source setting. Cross-source point cloud data corresponds to point sets captured by depth sensors with different accuracies or from different distances and perspectives. We address…

Cited by 7PDFcodeScholar
2023

DPP-Based Client Selection for Federated Learning with NON-IID DATA

ICASSP 2023accepted

This paper proposes a client selection (CS) method to tackle the communication bottleneck of federated learning (FL) while concurrently coping with FL’s data heterogeneity issue. Specifically, we first analyze the effect of CS in FL and show that FL training can be accelerated by adequately choosing…

Cited by 0SourceScholar
2023

METEOR: A Dense, Heterogeneous, and Unstructured Traffic Dataset with Rare Behaviors

ICRA 2023poster

We present a new traffic dataset, Meteor, which captures traffic patterns and multi-agent driving behaviors in unstructured scenarios. Meteor consists of more than 1000 one-minute videos, over 2 million annotated frames with bounding boxes and GPS trajectories for 16 unique agent categories, and mor…

Cited by 15SourceScholar
2023

Small-shot Multi-modal Distillation for Vision-based Autonomous Steering

ICRA 2023poster

In this paper, we propose a novel learning framework for autonomous systems that uses a small amount of “auxiliary information” that complements the learning of the main modality, called “small-shot auxiliary modality distillation network (AMD-S-Net)”. The AMD-S-Net contains a two-stream framework d…

Cited by 1SourceScholar
2022

FAR: Fourier Aerial Video Recognition

ECCV 2022poster

"We present a method, Fourier Activity Recognition (FAR), for UAV video activity recognition. Our formulation uses a novel Fourier object disentanglement method to innately separate out the human agent (which is typically small) from the background. Our disentanglement technique operates in the freq…

2019

Fully Learnable Group Convolution for Acceleration of Deep Neural Networks

CVPR 2019poster

Benefitted from its great success on many tasks, deep learning is increasingly used on low-computational-cost devices, e.g. smartphone, embedded devices, etc. To reduce the high computational and memory cost, in this work, we propose a fully learnable group convolution module (FLGC for short) which…

Cited by 94PDFScholar
2019

Spatially Adaptive Losses for Video Super-resolution with GANs

ICASSP 2019accepted

Deep Learning techniques and more specifically Generative Adversarial Networks (GANs) have recently been used for solving the video super-resolution (VSR) problem. In some of the published works, feature-based perceptual losses have also been used, resulting in promising results. While there has bee…

Cited by 0SourceScholar