← Search

Yan Pan

12 accepted papers

2026

GRIP: Latent Field-Guided Graph Policy for Budget-Constrained Multi-Agent Routing

AAAI 2026technical

Subset selection under budget constraints is critical in applications like multi-robot patrolling, crime deterrence, and targeted marketing, where multiple agents must jointly select targets and plan feasible routes. We formalize this challenge as Multi-Subset Selection with Budget-Constrained Routi

Cited by 0SourcePDFScholar
2026

Sequential Probabilistic Descriptor via Uncertainty-Aware Multi-Modal Fusion for Safety-Critical Place Recognition

RA-L 2026

The success of loop closure and global map consistency largely relies on robust place recognition, making it a safety-critical capability for mobile robots in autonomous navigation. However, most existing approaches prioritize recognition accuracy while overlooking uncertainty estimation, which may

Cited by 0SourceScholar
2025

Cool-Fusion: Fuse Large Language Models without Training

ACL 2025long

We focus on the problem of fusing two or more heterogeneous large language models (LLMs) to leverage their complementary strengths. One of the challenges of model fusion is high computational load, specifically in fine-tuning or aligning vocabularies. To address this, we propose Cool-Fusion, a simpl…

2025

On the Zero-shot Adversarial Robustness of Vision-Language Models: A Truly Zero-shot and Training-free Approach

CVPR 2025poster

Pre-trained Vision-Language Models (VLMs) like CLIP, have demonstrated strong zero-shot generalization capabilities. Despite their effectiveness on various downstream tasks, they remain vulnerable to adversarial samples. Existing methods fine-tune VLMs to improve their performance via performing adv…

Cited by 0SourcePDFScholar
2025

Thinking Before You Speak: A Proactive Test-time Scaling Approach

EMNLP 2025

Large Language Models (LLMs) often exhibit deficiencies with complex reasoning tasks, such as maths, which we attribute to the discrepancy between human reasoning patterns and those presented in the LLMs’ training data. When dealing with complex problems, humans tend to think carefully before expres

Cited by 0SourcePDFScholar
2024

MimicDiffusion: Purifying Adversarial Perturbation via Mimicking Clean Diffusion Model

CVPR 2024poster

Deep neural networks (DNNs) are vulnerable to adversarial perturbation where an imperceptible perturbation is added to the image that can fool the DNNs. Diffusion-based adversarial purification uses the diffusion model to generate a clean image against such adversarial attacks. Unfortunately the gen…

2023

Deep Hashing With Minimal-Distance-Separated Hash Centers

CVPR 2023poster

Deep hashing is an appealing approach for large-scale image retrieval. Most existing supervised deep hashing methods learn hash functions using pairwise or triple image similarities in randomly sampled mini-batches. They suffer from low training efficiency, insufficient coverage of data distribution…

Cited by 47SourcePDFScholar
2023

MotionBEV: Attention-Aware Online LiDAR Moving Object Segmentation With Bird's Eye View Based Appearance and Motion Features

RA-L 2023

Identifying moving objects is an essential capability for autonomous systems, as it provides critical information for pose estimation, navigation, collision avoidance, and static map construction. In this letter, we present MotionBEV, a fast and accurate framework for LiDAR moving object segmentatio

Cited by 26SourcecodeScholar
2022

Expressive Talking Head Generation With Granular Audio-Visual Control

CVPR 2022poster

Generating expressive talking heads is essential for creating virtual humans. However, existing one- or few-shot methods focus on lip-sync and head motion, ignoring the emotional expressions that make talking faces realistic. In this paper, we propose the Granularly Controlled Audio-Visual Talking H…

Cited by 148PDFScholar
2021

3DCaricShop: A Dataset and a Baseline Method for Single-View 3D Caricature Face Reconstruction

CVPR 2021poster

Caricature is an artistic representation that deliberately exaggerates the distinctive features of a human face to convey humor or sarcasm. However, reconstructing a 3D caricature from a 2D caricature image remains a challenging task, mostly due to the lack of data. We propose to fill this gap by in…

Cited by 27PDFScholar
2015

Simultaneous Feature Learning and Hash Coding With Deep Neural Networks

CVPR 2015poster

Similarity-preserving hashing is a widely-used method for nearest neighbour search in large-scale image retrieval tasks. For most existing hashing methods, an image is first encoded as a vector of hand-engineering visual features, followed by another separate projection or quantization step that gen…

Cited by 1028SourcePDFScholar