← Search

MingRui Lao

8 accepted papers

2026

PAUL: Uncertainty-Guided Partition and Augmentation for Robust Cross-View Geo-Localization under Noisy Correspondence

CVPR 2026

Cross-view geo-localization is a critical task for UAV navigation, event detection, and aerial surveying, which establish correspondence between drone-captured and satellite imagery. Most existing approaches embed cross-view data into a joint feature space to maximize similarity between paired image

Cited by 0SourceScholar
2026

Rep Deep & Machine Learning: Exemplar-Free Continual Video Action Recognition via Slow-Fast Collaborative Learning

AAAI 2026technical

In real-world applications, video action recognition models must continuously learn new action categories while retaining previously acquired knowledge. However, most existing approaches rely on storing historical data for replay, which introduces storage burdens and raises data privacy concerns. To

Cited by 0SourcePDFScholar
2025

Generalization-Preserved Learning: Closing the Backdoor to Catastrophic Forgetting in Continual Deepfake Detection

ICCV 2025poster

Existing continual deepfake detection methods typically treat stability (retaining previously learned forgery knowl- edge) and plasticity (adapting to novel forgeries) as con- flicting properties, emphasizing an inherent trade-off be- tween them, while regarding generalization to unseen forg- eries…

Cited by 0SourcePDFScholar
2025

Multi-Modal Entities Matter: Benchmarking Multi-Modal Entity Alignment

COLING 2025main

Multi-modal entity alignment (MMEA) is a long-standing task that aims to discover identical entities between different multi-modal knowledge graphs (MMKGs). However, most of the existing MMEA datasets consider the multi-modal data as the attributes of textual entities, while neglecting the correlati…

Cited by 0SourcePDFScholar
2024

Boosting Adversarial Robustness Distillation Via Hybrid Decomposed Knowledge

ICASSP 2024accepted

Adversarial Robust Distillation (ARD) has emerged as a potent defense mechanism tailored to small models against adversarial threats. However, mainstream ARD methods typically exploit teachers’ response as the transferred knowledge, while neglecting the analysis of involved target-related knowledge…

Cited by 0SourceScholar
2024

Language Without Borders: A Dataset and Benchmark for Code-Switching Lip Reading

NeurIPS 2024poster

Lip reading aims at transforming the videos of continuous lip movement into textual contents, and has achieved significant progress over the past decade. It serves as a critical yet practical assistance for speech-impaired individuals, with more practicability than speech recognition in noisy enviro…

2023

COCA: COllaborative CAusal Regularization for Audio-Visual Question Answering

AAAI 2023technical

Audio-Visual Question Answering (AVQA) is a sophisticated QA task, which aims at answering textual questions over given video-audio pairs with comprehensive multimodal reasoning. Through detailed causal-graph analyses and careful inspections of their learning processes, we reveal that AVQA models ar…

Cited by 21SourcePDFScholar
2022

VQA-BC: Robust Visual Question Answering Via Bidirectional Chaining

ICASSP 2022accepted

Current VQA models are suffering from the problem of overdependence on language bias, which severely reduces their robustness in real-world scenarios. In this paper, we analyze VQA models from the view of forward/backward chaining in the inference engine, and propose to enhance their robustness via…

Cited by 0SourceScholar