← Search

Lijun He

7 accepted papers

2026

HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly Detection

AAAI 2026technical

Video Anomaly Detection (VAD) aims to locate events that deviate from normal patterns in videos. Traditional approaches often rely on extensive labeled data and incur high computational costs. Recent tuning-free methods based on Multimodal Large Language Models (MLLMs) offer a promising alternative

Cited by 0SourcePDFScholar
2026

Invisible Triggers, Visible Threats! Road-Style Adversarial Creation Attack for Visual 3D Detection in Autonomous Driving

AAAI 2026technical

Modern autonomous driving (AD) systems leverage 3D object detection to perceive foreground objects in 3D environments for subsequent prediction and planning. Visual 3D detection based on RGB cameras provides a cost-effective solution compared to the LiDAR paradigm. While achieving promising detectio

Cited by 0SourcePDFScholar
2026

Steering and Rectifying Latent representation manifolds in Frozen Multi-modal LLMs for Video Anomaly Detection

ICLR 2026poster

Video anomaly detection (VAD) aims to identify abnormal events in videos. Traditional VAD methods generally suffer from the high costs of labeled data and full training, thus some recent works have explored leveraging frozen multi-modal large language models (MLLMs) in a tuning-free manner to perfor…

Cited by 0SourceScholar
2026

Unleashing the Representational Power of Fourier Shapes for Attacking Infrared Object Detection

ICML 2026poster

Infrared object detection is crucial for perception in autonomous driving and surveillance but remains vulnerable to physical adversarial attacks. Unlike in the RGB domain, where attacks rely on color texture, infrared attacks must manipulate thermal signatures, making the geometry shape of heat-blo…

Cited by 0SourceScholar
2026

What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNs

CVPR 2026

Split DNNs enable edge devices by offloading intensive computation to a cloud server, but this paradigm exposes privacy vulnerabilities, as the intermediate features can be exploited to reconstruct the private inputs via Feature Inversion Attack (FIA). Existing FIA methods often produce limited reco

Cited by 0SourceScholar
2025

A Singing Melody Extraction Network Via Self-Distillation and Multi-Level Supervision

ICASSP 2025accepted

Extracting singing melody from polyphonic music is an important topic in the field of music information retrieval. In this paper, we propose a singing melody extraction network consisting of five stacked multi-scale feature time-frequency aggregation (MF-TFA) modules. In the same network, deeper lay…

Cited by 0SourceScholar
2024

Fine-grained Dynamic Network for Generic Event Boundary Detection

ECCV 2024poster

"Generic event boundary detection (GEBD) aims at pinpointing event boundaries naturally perceived by humans, playing a crucial role in understanding long-form videos. Given the diverse nature of generic boundaries, spanning different video appearances, objects, and actions, this task remains challen…