← Search

Xingjian Wang

9 accepted papers

2026

Distributed Bearing-Only Formation Maneuvering Control for Quadrotors Without Global Reference Frame

RA-L 2026

Most existing bearing-only formation control methods required that the relative bearings among neighboring agents are measured under a well-known global reference frame for each individual. To remove such constraint, this paper novelly introduces a distributed formation control scheme for quadrotors

Cited by 0SourceScholar
2026

Distributed Bearing-Only Formation Maneuvering Control for Quadrotors without Global Reference Frame

ICRA 2026poster

Most existing bearing-only formation control methods required that the relative bearings among neighboring agents are measured under a well-known global reference frame for each individual. To remove such constraint, this paper novelly introduces a distributed formation control scheme for quadrotors…

Cited by 0SourceScholar
2026

FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation

AAAI 2026technical

Recent unified models have demonstrated that the reasoning capacity of Multimodal Large Language Models (MLLMs) can be leveraged to facilitate diffusion-based image generation with impressive flexibility and performance. However, approaches that rely heavily on MLLMs for high-level semantic encoding

Cited by 0SourcePDFScholar
2026

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

ICLR 2026poster

Understanding long videos requires Multimodal Large Language Models (MLLMs) to grasp multi-timescale information, often organized in hierarchies. However, current long-video understanding benchmarks either overlook multi-timescale design or distribute questions targeting different timescales across…

Cited by 0SourcecodeScholar
2026

Synthetic Curriculum Reinforces Compositional Text-to-Image Generation

CVPR 2026

Text-to-Image (T2I) generation has long been an open problem, with compositional synthesis remaining particularly challenging. This task requires accurate rendering of complex scenes containing multiple objects that exhibit diverse attributes as well as intricate spatial and semantic relationships,

Cited by 0SourceScholar
2025

Debiasing Trace Guidance: Top-down Trace Distillation and Bottom-up Velocity Alignment for Unsupervised Anomaly Detection

ICCV 2025poster

The leak of anomalous information from input condition poses a great challenge to reconstruction-based anomaly detection. Recent diffusion-based methods respond to this issue by suppressing anomaly information for condition injection or in-sampling inversion. However, since they treat conditions as…

Cited by 0SourcePDFScholar
2025

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation

NeurIPS 2025spotlight

Recent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities. Existing benchmarks are often domain-specific and thus cannot fully capture an LLM’s general reasoning potential. To address this limit…

Cited by 0SourcecodeScholar
2025

Lifting Scheme-Based Implicit Disentanglement of Emotion-Related Facial Dynamics in the Wild

AAAI 2025technical

In-the-wild dynamic facial expression recognition (DFER) encounters a significant challenge in recognizing emotion-related expressions, which are often temporally and spatially diluted by emotion-irrelevant expressions and global context. Most prior DFER methods directly utilize coupled spatiotempor…

2023

Low in Resolution, High in Precision: UAV Detection with Super-Resolution and Motion Information Extraction

ICASSP 2023accepted

The rapid development of unmanned aerial vehicle (UAV) market presents potential threats to public security and personal privacy, and the vision sensors are widely deployed to detect the invasive UAVs because of the intuitivity and accessibility of the video. However, the small pixel area and weak m…

Cited by 0SourceScholar