← Search

Pengyu Li

13 accepted papers

2026

Progressive Multi-modal Knowledge Distillation for Multi-spectral Object Re-identification

AAAI 2026technical

In the field of multi-spectral object re-identification (ReID), multi-modal knowledge and modal-specific knowledge exhibit complementary advantages when handling hard samples, but existing methods rarely integrate this collaborative information. Knowledge distillation is a direct approach for trans

Cited by 0SourcePDFScholar
2024

Neural Collapse in Multi-label Learning with Pick-all-label Loss

ICML 2024poster

We study deep neural networks for the multi-label classification (MLab) task through the lens of neural collapse (NC). Previous works have been restricted to the multi-class classification setting and discovered a prevalent NC phenomenon comprising of the following properties for the last-layer feat…

2023

FastInst: A Simple Query-Based Model for Real-Time Instance Segmentation

CVPR 2023poster

Recent attention in instance segmentation has focused on query-based models. Despite being non-maximum suppression (NMS)-free and end-to-end, the superiority of these models on high-accuracy real-time benchmarks has not been well demonstrated. In this paper, we show the strong potential of query-bas…

2023

Learning Polysemantic Spoof Trace: A Multi-Modal Disentanglement Network for Face Anti-spoofing

AAAI 2023technical

Along with the widespread use of face recognition systems, their vulnerability has become highlighted. While existing face anti-spoofing methods can be generalized between attack types, generic solutions are still challenging due to the diversity of spoof characteristics. Recently, the spoof trace d…

Cited by 4SourcePDFScholar
2023

Longshortnet: Exploring Temporal and Semantic Features Fusion In Streaming Perception

ICASSP 2023accepted

Streaming perception is a fundamental task in autonomous driving that requires a careful balance between the latency and accuracy of the autopilot system. However, current methods for streaming perception are limited as they rely only on the current and adjacent two frames to learn movement patterns…

Cited by 0SourceScholar
2023

Optimal Proposal Learning for Deployable End-to-End Pedestrian Detection

CVPR 2023poster

End-to-end pedestrian detection focuses on training a pedestrian detection model via discarding the Non-Maximum Suppression (NMS) post-processing. Though a few methods have been explored, most of them still suffer from longer training time and more complex deployment, which cannot be deployed in the…

Cited by 19SourcePDFScholar
2022

Dense Learning Based Semi-Supervised Object Detection

CVPR 2022poster

The ultimate goal of semi-supervised object detection (SSOD) is to facilitate the utilization and deployment of detectors in actual applications with the help of a large amount of unlabeled data. Although a few works have proposed various self-training-based methods or consistency-regularization-bas…

Cited by 87PDFcodeScholar
2021

Adversarial Pose Regression Network for Pose-Invariant Face Recognitions

AAAI 2021technical

Face recognition has achieved significant progress in recent years. However, the large pose variation between face images remains a challenge in face recognition. We observe that the pose variation in the hidden feature maps is one of the most critical factors to hinder the representations from bein…

2021

Variational Attention: Propagating Domain-Specific Knowledge for Multi-Domain Learning in Crowd Counting

ICCV 2021poster

In crowd counting, due to the problem of laborious labelling, it is perceived intractability of collecting a new large-scale dataset which has plentiful images with large diversity in density, scene, etc. Thus, for learning a general model, training with data from multiple different datasets might b…

Cited by 57PDFcodeScholar
2021

VirFace: Enhancing Face Recognition via Unlabeled Shallow Data

CVPR 2021poster

Recently, exploiting the effect of the unlabeled data for face recognition attracts increasing attention. However, there are still few works considering the situation that the unlabeled data is shallow which widely exists in real-world scenarios. The existing semi-supervised face recognition methods…

Cited by 22PDFScholar
2021

Virtual Fully-Connected Layer: Training a Large-Scale Face Recognition Dataset With Limited Computational Resources

CVPR 2021poster

Recently, deep face recognition has achieved significant progress because of Convolutional Neural Networks (CNNs) and large-scale datasets. However, training CNNs on a large-scale face recognition dataset with limited computational resources is still a challenge. This is because the classification p…

Cited by 29PDFcodeScholar