← Search

Hui Miao

4 accepted papers

2025

Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt Learning

AAAI 2025technical

With the malicious use and dissemination of multi-modal deepfake videos, researchers start to investigate multi-modal deepfake detection. Unfortunately, most of the existing methods tune all the parameters of the deep network with limited speech video datasets and are trained under coarse-grained co…

Cited by 0SourcePDFScholar
2024

Distilling Vision-Language Models on Millions of Videos

CVPR 2024poster

The recent advance in vision-language models is largely attributed to the abundance of image-text data. We aim to replicate this success for video-language models but there simply is not enough human-curated video-text data available. We thus resort to fine-tuning a video-language model from a stron…

Cited by 18SourcePDFScholar
2021

Robust 2D/3D Vehicle Parsing in Arbitrary Camera Views for CVIS

ICCV 2021poster

We present a novel approach to robustly detect and perceive vehicles in different camera views as part of a cooperative vehicle-infrastructure system (CVIS). Our formulation is designed for arbitrary camera views and makes no assumptions about intrinsic or extrinsic parameters. First, to deal with m…

Cited by 3PDFcodeScholar
2020

3D Part Guided Image Editing for Fine-Grained Object Understanding

CVPR 2020poster

Holistically understanding an object with its 3D movable parts is essential for visual models of a robot to interact with the world. For example, only by understanding many possible part dynamics of other vehicles (e.g., door or trunk opening, taillight blinking for changing lane), a self-driving ve…

Cited by 14PDFcodeScholar