← Search

Boon Poh Ng

8 accepted papers

2026

From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge

AAAI 2026technical

Large-scale Video Foundation Models (VFMs) have significantly advanced various video-related tasks, either through task-specific models or Multi-modal Large Language Models (MLLMs). However, the open accessibility of VFMs also introduces critical security risks, as adversaries can exploit full knowl

Cited by 0SourcePDFScholar
2025

Vid-Group: Temporal Video Grounding Pretraining from Unlabeled Videos in the Wild

ICCV 2025poster

Given a natural language query, temporal video grounding aims to localize the described temporal moment in an untrimmed video. A major challenge of this task is its heavy dependence on labor-intensive annotations for training. Unlike existing works that directly train models on manually curated data…

2024

E3M: Zero-Shot Spatio-Temporal Video Grounding with Expectation-Maximization Multimodal Modulation

ECCV 2024oral

"Spatio-temporal video grounding aims to localize the spatio-temporal tube in a video according to the given language query. To eliminate the annotation costs, we make a first exploration to tackle spatio-temporal video grounding in a zero-shot manner. Our method dispenses with the need for any trai…

2024

Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video Grounding

AAAI 2024technical

This paper for the first time leverages multi-modal videos for weakly-supervised temporal video grounding. As labeling the video moment is labor-intensive and subjective, the weakly-supervised approaches have gained increasing attention in recent years. However, these approaches could inherently com…

Cited by 12SourcePDFScholar
2024

Omnipotent Distillation with LLMs for Weakly-Supervised Natural Language Video Localization: When Divergence Meets Consistency

AAAI 2024technical

Natural language video localization plays a pivotal role in video understanding, and leveraging weakly-labeled data is considered a promising approach to circumvent the laborintensive process of manual annotations. However, this approach encounters two significant challenges: 1) limited input distri…

Cited by 9SourcePDFScholar
2023

Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event Localization

AAAI 2023technical

This paper for the first time explores audio-visual event localization in an unsupervised manner. Previous methods tackle this problem in a supervised setting and require segment-level or video-level event category ground-truth to train the model. However, building large-scale multi-modality dataset…

Cited by 9SourcePDFScholar
2022

An Adaptive Orientational Beamforming Technique for Narrowband Interference Rejection

ICASSP 2022accepted

In this paper, we investigate and extend the linearly constrained minimum variance (LCMV) algorithm for conventional wideband beamforming system to the recently proposed orientational beamforming (OBF) system. An orientational LCMV (O-LCMV) algorithm is proposed. It is constructed on the orientation…

Cited by 0SourceScholar
2021

Learning Sparsifying Transforms for Image Reconstruction in Electrical Impedance Tomography

ICASSP 2021accepted

Electrical Impedance Tomography (EIT) is a fast and non-invasive imaging technology that reconstructs the internal electrical properties of a subject. However, its functionality is limited by low spatial resolution arising from an ill-posed and ill-conditioned inverse problem. Several sparsity-promo…

Cited by 0SourceScholar