← Search

Xian Zhong

18 accepted papers

2026

Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order Transfer

AAAI 2026technical

Action recognition using uncrewed aerial vehicles (UAVs) faces unique challenges due to substantial view variations along the vertical spatial axis. Unlike ground-based scenarios, UAVs capture actions from diverse altitudes, resulting in pronounced appearance discrepancies and reduced recognition ro

Cited by 0SourcePDFScholar
2026

Evolving Quantitative Reasoning through Self-Play in Digital Twin Markets

ICML 2026poster

Large Language Models (LLMs) exhibit strong capabilities in high-level semantic understanding and strategic planning, yet they suffer from persistent quantitative failure modes, such as imprecise computation and the illusion of quantitative coherence, which limit their reliability in high-stakes dec…

Cited by 0SourceScholar
2025

Agent Trading Arena: A Study on Numerical Understanding in LLM-Based Agents

EMNLP 2025

Large language models (LLMs) have demonstrated remarkable capabilities in natural language tasks, yet their performance in dynamic, real-world financial environments remains underexplored. Existing approaches are confined to historical backtesting, where trading actions cannot influence market price

2025

LGNet: Linear Graph Representation for Efficient Cold-Start Recommendations

ICASSP 2025accepted

Graph Convolutional Networks (GCNs) demonstrate significant potential in recommendation systems but face difficulties with the cold-start problem, especially in integrating new nodes during inference. The typical solution leverages meta-learning for few-shot learning, though it often fails to fully…

Cited by 0SourceScholar
2025

SOTA: Spike-Navigated Optimal TrAnsport Saliency Region Detection in Composite-bias Videos

IJCAI 2025

Existing saliency detection methods struggle in real-world scenarios due to motion blur and occlusions. In contrast, spike cameras, with their high temporal resolution, significantly enhance visual saliency maps. However, the composite noise inherent to spike camera imaging introduces discontinuitie

2025

STAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural Networks

CVPR 2025poster

Spiking Neural Networks (SNNs) have gained significant attention due to their biological plausibility and energy efficiency, making them promising alternatives to Artificial Neural Networks (ANNs). However, the performance gap between SNNs and ANNs remains a substantial challenge hindering the wides…

Cited by 0SourcePDFScholar
2025

StoryLLaVA: Enhancing Visual Storytelling with Multi-Modal Large Language Models

COLING 2025main

The rapid development of multimodal large language models (MLLMs) has positioned visual storytelling as a crucial area in content creation. However, existing models often struggle to maintain temporal, spatial, and narrative coherence across image sequences, and they frequently lack the depth and en…

Cited by 3SourcePDFScholar
2025

Synergistic Integration of Cross-Spatial Learning for Lightweight Crack Detection

ICASSP 2025accepted

Efficient crack segmentation is crucial for engineering surface inspection, especially on edge devices where both accuracy and computational efficiency are essential. To address the challenges posed by crack directionality and blurred edges while enhancing performance, we propose a lightweight segme…

Cited by 0SourceScholar
2023

Background Disturbance Mitigation for Video Captioning Via Entity-Action Relocation

ICASSP 2023accepted

Video captioning aims to generate sentences to accurately describe the video content, in which video background plays the role of prompts. State-of-the-art methods tend to explore richer video representations adequately, fusing with language to improve caption quality, which has shown great success.…

Cited by 0SourceScholar
2023

Background-Weakening Consistency Regularization for Semi-Supervised Video Action Detection

ICASSP 2023accepted

Consistency-based techniques have produced state-of-the-art results in semi-supervised action detection. When the model false detects the dynamic information in the background as an action, spatio-temporal consistency calculations can hardly reflect this false detection result. We consider weakening…

Cited by 0SourceScholar
2023

Bat: Bi-Alignment Based On Transformation in Multi-Target Domain Adaptation for Semantic Segmentation

ICASSP 2023accepted

While enlightening progress has been made recently in single-target domain adaptive semantic segmentation (ST-DASS), the multi-peak distributed multi-target domain cannot be directly aligned well with the single-peak distributed source domain. As a result, it is impossible for existing methods to ha…

Cited by 0SourceScholar
2023

Good Is Bad: Causality Inspired Cloth-Debiasing for Cloth-Changing Person Re-Identification

CVPR 2023poster

Entangled representation of clothing and identity (ID)-intrinsic clues are potentially concomitant in conventional person Re-IDentification (ReID). Nevertheless, eliminating the negative impact of clothing on ID remains challenging due to the lack of theory and the difficulty of isolating the exact…

2023

Neighborhood Information-Based Label Refinement for Person Re-Identification with Label Noise

ICASSP 2023accepted

The existing excellent person re-identification (Re-ID) model is still affected by the samples with the incorrect labels. It is difficult to accurately annotate person images in the real scene, resulting in label noise. To avoid fitting to the noisy labels, a common solution in Re-ID is to replace t…

Cited by 0SourceScholar
2023

Refined Semantic Enhancement towards Frequency Diffusion for Video Captioning

AAAI 2023technical

Video captioning aims to generate natural language sentences that describe the given video accurately. Existing methods obtain favorable generation by exploring richer visual representations in encode phase or improving the decoding ability. However, the long-tailed problem hinders these attempts at…

2022

Rainy WCity: A Real Rainfall Dataset with Diverse Conditions for Semantic Driving Scene Understanding

IJCAI 2022poster

Scene understanding in adverse weather conditions (e.g. rainy and foggy days) has drawn increasing attention, arising some specific benchmarks and algorithms. However, scene segmentation under rainy weather is still challenging and under-explored due to the following limitations on the datasets and…

Cited by 32SourcePDFScholar
2022

VCD: View-Constraint Disentanglement for Action Recognition

ICASSP 2022accepted

Action recognition is a hot topic in computer vision due to its wide range of applications in urban surveillance. Although some methods are more advanced from an invariant view perspective, those approaches do not perform well for the viewpoint change. To address this issue, one possible solution is…

Cited by 0SourceScholar
2021

Part-Aligned Network with Background for Misaligned Person Search

ICASSP 2021accepted

Person search is a significant computer vision task that requires addressing person detection and re-identification simultaneously. Body parts are frequently misaligned due to variation poses, occlusions, and partial missing, leading to the unsatisfied results of person search. Existing methods usua…

Cited by 0SourceScholar
2020

Multi-Scale Residual Network for Image Classification

ICASSP 2020accepted

Multi-scale approach representing image objects at various levels-of-details has been applied to various computer vision tasks. Existing image classification approaches place more emphasis on multi-scale convolution kernels, and overlook multi-scale feature maps. As such, some shallower information…

Cited by 0SourceScholar