← Search

Mengzhu Li

3 accepted papers

2026

Beyond Instance-Level Self-Supervision in 3D Multi-Modal Medical Imaging

ICML 2026poster

Self-supervised pre-training methods in medical imaging typically treat each individual as an isolated instance, learning representations through augmentation-based objectives or masked reconstruction. They often do not adequately capitalize on a key characteristic of physiological features: anatomi…

Cited by 0SourceScholar
2022

Transtl: Spatial-Temporal Localization Transformer for Multi-Label Video Classification

ICASSP 2022accepted

Multi-label video classification (MLVC) is a long-standing and challenging research problem in video signal analysis. Generally, there exist many complex action labels in real-world videos and these actions are with inherent dependencies at both spatial and temporal domains. Motivated by this observ…

Cited by 0SourceScholar
2022

W-ART: Action Relation Transformer for Weakly-Supervised Temporal Action Localization

ICASSP 2022accepted

Weakly-supervised temporal action localization (WTAL) is a long-standing and challenging research problem in video signal analysis. It is to localize the action segments in the video given only video-level labels. The key to this task is understanding how the diverse actions interact. In this paper,…

Cited by 0SourceScholar