← Search

Meng Hwa Er

7 accepted papers

2025

Vid-Group: Temporal Video Grounding Pretraining from Unlabeled Videos in the Wild

ICCV 2025poster

Given a natural language query, temporal video grounding aims to localize the described temporal moment in an untrimmed video. A major challenge of this task is its heavy dependence on labor-intensive annotations for training. Unlike existing works that directly train models on manually curated data…

2024

Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video Grounding

AAAI 2024technical

This paper for the first time leverages multi-modal videos for weakly-supervised temporal video grounding. As labeling the video moment is labor-intensive and subjective, the weakly-supervised approaches have gained increasing attention in recent years. However, these approaches could inherently com…

Cited by 12SourcePDFScholar
2024

Omnipotent Distillation with LLMs for Weakly-Supervised Natural Language Video Localization: When Divergence Meets Consistency

AAAI 2024technical

Natural language video localization plays a pivotal role in video understanding, and leveraging weakly-labeled data is considered a promising approach to circumvent the laborintensive process of manual annotations. However, this approach encounters two significant challenges: 1) limited input distri…

Cited by 9SourcePDFScholar
2023

Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event Localization

AAAI 2023technical

This paper for the first time explores audio-visual event localization in an unsupervised manner. Previous methods tackle this problem in a supervised setting and require segment-level or video-level event category ground-truth to train the model. However, building large-scale multi-modality dataset…

Cited by 9SourcePDFScholar
2022

An Adaptive Orientational Beamforming Technique for Narrowband Interference Rejection

ICASSP 2022accepted

In this paper, we investigate and extend the linearly constrained minimum variance (LCMV) algorithm for conventional wideband beamforming system to the recently proposed orientational beamforming (OBF) system. An orientational LCMV (O-LCMV) algorithm is proposed. It is constructed on the orientation…

Cited by 0SourceScholar
2021

Skeleton Cloud Colorization for Unsupervised 3D Action Representation Learning

ICCV 2021poster

Skeleton-based human action recognition has attracted increasing attention in recent years. However, most of the existing works focus on supervised learning which requiring a large number of annotated action sequences that are often expensive to collect. We investigate unsupervised representation le…

Cited by 123PDFScholar
2020

Collaborative Learning of Gesture Recognition and 3D Hand Pose Estimation with Multi-Order Feature Analysis

ECCV 2020poster

Gesture recognition and 3D hand pose estimation are two highly correlated tasks, yet they are often handled separately. In this paper, we present a novel collaborative learning network for joint gesture recognition and 3D hand pose estimation. The proposed network exploits joint-aware features that…

Cited by 57SourcePDFScholar