← Search

Ji Gan

6 accepted papers

2026

Linguistic Relative Policy Optimization for Video Anomaly Reasoning

ICML 2026poster

Video anomaly detection (VAD) with multimodal large language models has shown strong potential, yet most existing methods still depend on large-scale annotations or expert-designed priors, limiting their ability to acquire anomaly knowledge with as little human intervention as possible. To address t…

Cited by 0SourceScholar
2025

A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly Understanding

NeurIPS 2025poster

While unmanned aerial vehicles (UAVs) offer wide-area, high-altitude coverage for anomaly detection, they face challenges such as dynamic viewpoints, scale variations, and complex scenes. Existing datasets and methods, mainly designed for fixed ground-level views, struggle to adapt to these conditio…

Cited by 0SourcecodeScholar
2025

Structure-Aware Handwritten Text Recognition via Graph-Enhanced Cross-Modal Mutual Learning

IJCAI 2025

Existing handwriting recognition methods only focus on learning visual patterns by modeling low-level relationships of adjacent pixels, while overlooking the intrinsic geometric structures of characters. In this paper, we propose a novel graph-enhanced cross-modal mutual learning network GCM to full

Cited by 0SourcePDFScholar
2024

Beyond Euclidean: Dual-Space Representation Learning for Weakly Supervised Video Violence Detection

NeurIPS 2024poster

While numerous Video Violence Detection (VVD) methods have focused on representation learning in Euclidean space, they struggle to learn sufficiently discriminative features, leading to weaknesses in recognizing normal events that are visually similar to violent events (i.e., ambiguous violence). In…

Cited by 3SourcePDFScholar
2024

Structure-Aware in-Air Handwritten Text Recognition with Graph-Guided Cross-Modality Translator

ICASSP 2024accepted

In-air handwriting as a new human-computer interaction way plays an important role in many virtual/mixed-reality applications. Existing methods for in-air handwritten text recognition (IAHTR) typically directly process handwriting trajectories with deep neural networks. However, those methods all si…

Cited by 0SourceScholar
2021

HiGAN: Handwriting Imitation Conditioned on Arbitrary-Length Texts and Disentangled Styles

AAAI 2021technical

Given limited handwriting scripts, humans can easily visualize (or imagine) what the handwritten words/texts would look like with other arbitrary textual contents. Moreover, a person also is able to imitate the handwriting styles of provided reference samples. Humans can do such hallucinations, perh…