← Search

Peijun Bao

7 accepted papers

2026

ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos

CVPR 2026

Temporal forgery localization aims to temporally identify manipulated segments in videos. Most existing benchmarks focus on appearance-level forgeries, such as face swapping and object removal. However, recent advances in video generation have driven the emergence of activity-level forgeries that mo

Cited by 0SourcecodeScholar
2025

Vid-Group: Temporal Video Grounding Pretraining from Unlabeled Videos in the Wild

ICCV 2025poster

Given a natural language query, temporal video grounding aims to localize the described temporal moment in an untrimmed video. A major challenge of this task is its heavy dependence on labor-intensive annotations for training. Unlike existing works that directly train models on manually curated data…

2024

Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video Grounding

AAAI 2024technical

This paper for the first time leverages multi-modal videos for weakly-supervised temporal video grounding. As labeling the video moment is labor-intensive and subjective, the weakly-supervised approaches have gained increasing attention in recent years. However, these approaches could inherently com…

Cited by 12SourcePDFScholar
2024

Omnipotent Distillation with LLMs for Weakly-Supervised Natural Language Video Localization: When Divergence Meets Consistency

AAAI 2024technical

Natural language video localization plays a pivotal role in video understanding, and leveraging weakly-labeled data is considered a promising approach to circumvent the laborintensive process of manual annotations. However, this approach encounters two significant challenges: 1) limited input distri…

Cited by 9SourcePDFScholar
2023

Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event Localization

AAAI 2023technical

This paper for the first time explores audio-visual event localization in an unsupervised manner. Previous methods tackle this problem in a supervised setting and require segment-level or video-level event category ground-truth to train the model. However, building large-scale multi-modality dataset…

Cited by 9SourcePDFScholar
2021

Learning 3-D Human Pose Estimation from Catadioptric Videos

IJCAI 2021poster

3-D human pose estimation is a crucial step for understanding human actions. However, reliably capturing precise 3-D position of human joints is non-trivial and tedious. Current models often suffer from the scarcity of high-quality 3-D annotated training data. In this work, we explore a novel way of…

Cited by 4SourcePDFScholar