← Search

Jiaxu Leng

13 accepted papers

2026

Dynamic-Static Collaboration for Unsupervised Domain Adaptive Video-Based Visible-Infrared Person Re-Identification

AAAI 2026technical

Video-based visible-infrared person re-identification (VVI-ReID) aims to match pedestrian sequences across modalities for all-day surveillance. While supervised methods have shown progress, their dependence on large-scale cross-modal annotations limits scalability. We investigate the task of unsuper

Cited by 0SourcePDFScholar
2026

Hyperbolic Hierarchical Alignment for Video-Based Visible-Infrared Person Re-Identification

ICML 2026poster

Video-based visible-infrared person re-identification (VVI-ReID) aims to learn robust video-level representations under modality discrepancy. However, existing methods typically rely on Euclidean geometry, which is suboptimal for modeling the complex temporal dynamics within visible and infrared tra…

Cited by 0SourceScholar
2026

Learning to Watch: Active Video Anomaly Understanding via Interleaved Policy Optimization

ICML 2026poster

Video anomaly understanding (VAU) relies on sparse, context-dependent cues. However, existing passive paradigms suffer from observational aliasing, where static sampling fails to disambiguate semantically distinct events. To overcome this, we propose $Anom\text{-}\pi$, a closed-loop framework that r…

Cited by 0SourceScholar
2026

Linguistic Relative Policy Optimization for Video Anomaly Reasoning

ICML 2026poster

Video anomaly detection (VAD) with multimodal large language models has shown strong potential, yet most existing methods still depend on large-scale annotations or expert-designed priors, limiting their ability to acquire anomaly knowledge with as little human intervention as possible. To address t…

Cited by 0SourceScholar
2026

PortraitSR: Artist-Inspired Prior Learning for Progressive Face Super-Resolution

AAAI 2026technical

Face super-resolution (FSR) aims to reconstruct high-resolution (HR) face images from low-resolution (LR) inputs. While recent methods have advanced this task through architectural innovations and generative modeling, but they often leads to semantically inconsistent structures and unrealistic textu

Cited by 0SourcePDFScholar
2026

R2-LIO: Real-Time and Robust LiDAR-Inertial Odometry in Dynamic Environments

ICRA 2026poster

LiDAR-Inertial Odometry (LIO) is crucial for robot navigation and autonomous driving. Most existing methods rely on the assumption of a static environment, indiscriminately using all LiDAR measurements for localization. However, LiDAR data acquired in urban scenes often contain dynamic objects such …

Cited by 0Scholar
2026

Towards Trustworthy Video Anomaly Understanding: A Class-Guided Chain-of-Evaluation Metric and An Anomaly-focused Meta-Benchmark

ICML 2026poster

The trustworthiness of evaluation is critical to reliable model comparison and deployment in Video Anomaly Understanding (VAU). However, existing metrics are sensitive to expression styles and normal content, and this field lacks a diagnostic benchmark to validate metric validity and robustness. To …

Cited by 0SourceScholar
2025

A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly Understanding

NeurIPS 2025poster

While unmanned aerial vehicles (UAVs) offer wide-area, high-altitude coverage for anomaly detection, they face challenges such as dynamic viewpoints, scale variations, and complex scenes. Existing datasets and methods, mainly designed for fixed ground-level views, struggle to adapt to these conditio…

Cited by 0SourcecodeScholar
2025

Bidirectional Reference Image Quality Assessment via Content-Quality Correlation Modeling

ICASSP 2025accepted

The emphasis on no-reference image quality assessment has often overshadowed the significance of Full-Reference Image Quality Assessment (FR-IQA), which generally better reflects human contrastive perception mechanism. However, FRIQA presents challenges in obtaining content-aligned reference images.…

Cited by 0SourceScholar
2025

Structure-Aware Handwritten Text Recognition via Graph-Enhanced Cross-Modal Mutual Learning

IJCAI 2025

Existing handwriting recognition methods only focus on learning visual patterns by modeling low-level relationships of adjacent pixels, while overlooking the intrinsic geometric structures of characters. In this paper, we propose a novel graph-enhanced cross-modal mutual learning network GCM to full

Cited by 0SourcePDFScholar
2024

Beyond Euclidean: Dual-Space Representation Learning for Weakly Supervised Video Violence Detection

NeurIPS 2024poster

While numerous Video Violence Detection (VVD) methods have focused on representation learning in Euclidean space, they struggle to learn sufficiently discriminative features, leading to weaknesses in recognizing normal events that are visually similar to violent events (i.e., ambiguous violence). In…

Cited by 3SourcePDFScholar
2024

MGRL: Mutual-Guidance Representation Learning for Text-to-Image Person Retrieval

ICASSP 2024accepted

Text-to-image person retrieval aims to recognize target pedestrians based on specified text. Existing methods mainly obtain image and text features separately through distinct feature extractors, subsequently embedding them into a unified feature space and calculating their similarity. Despite great…

Cited by 0SourceScholar
2024

Structure-Aware in-Air Handwritten Text Recognition with Graph-Guided Cross-Modality Translator

ICASSP 2024accepted

In-air handwriting as a new human-computer interaction way plays an important role in many virtual/mixed-reality applications. Existing methods for in-air handwritten text recognition (IAHTR) typically directly process handwriting trajectories with deep neural networks. However, those methods all si…

Cited by 0SourceScholar