← Search

Yan Bai

12 accepted papers

2025

DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving

CVPR 2025highlight

Multimodal large language models (MLLMs) possess the ability to comprehend visual images or videos, and show impressive reasoning ability thanks to the vast amounts of pretrained knowledge, making them highly suitable for autonomous driving applications. Unlike the previous work, DriveGPT4-V1, which…

Cited by 0SourcePDFScholar
2024

Global-Local Collaborative Inference with LLM for Lidar-Based Open-Vocabulary Detection

ECCV 2024poster

"Open-Vocabulary Detection (OVD) is the task of detecting all interesting objects in a given scene without predefined object classes. Extensive work has been done to deal with the OVD for 2D RGB images, but the exploration of 3D OVD is still limited. Intuitively, lidar point clouds provide 3D inform…

2023

Decorate the Newcomers: Visual Domain Prompt for Continual Test Time Adaptation

AAAI 2023technical

Continual Test-Time Adaptation (CTTA) aims to adapt the source model to continually changing unlabeled target domains without access to the source data. Existing methods mainly focus on model-based adaptation in a self-training manner, such as predicting pseudo labels for new domain datasets. Since…

Cited by 100SourcePDFScholar
2023

Switchable Representation Learning Framework With Self-Compatibility

CVPR 2023poster

Real-world visual search systems involve deployments on multiple platforms with different computing and storage resources. Deploying a unified model that suits the minimal-constrain platforms leads to limited accuracy. It is expected to deploy models with different capacities adapting to the resourc…

Cited by 3SourcePDFScholar
2022

Neighborhood Consensus Contrastive Learning for Backward-Compatible Representation

AAAI 2022technical

In object re-identification (ReID), the development of deep learning techniques often involves model updates and deployment. It is unbearable to re-embedding and re-index with the system suspended when deploying new models. Therefore, backward-compatible representation is proposed to enable ``new''…

Cited by 8SourcePDFScholar
2021

Federated Learning for Non-IID Data via Unified Feature Learning and Optimization Objective Alignment

ICCV 2021poster

Federated Learning (FL) aims to establish a shared model across decentralized clients under the privacy-preserving constraint. Despite certain success, it is still challenging for FL to deal with non-IID (non-independent and identical distribution) client data, which is a general scenario in real-wo…

Cited by 100PDFScholar
2021

Person30K: A Dual-Meta Generalization Network for Person Re-Identification

CVPR 2021poster

Recently, person re-identification (ReID) has vastly benefited from the surging waves of data-driven methods. However, these methods are still not reliable enough for real-world deployments, due to the insufficient generalization capability of the models learned on existing benchmarks that have limi…

Cited by 78PDFScholar
2020

Disentangled Feature Learning Network for Vehicle Re-Identification

IJCAI 2020poster

Vehicle Re-Identification (ReID) has attracted lots of research efforts due to its great significance to the public security. In vehicle ReID, we aim to learn features that are powerful in discriminating subtle differences between vehicles which are visually similar, and also robust against differen…

Cited by 0SourcePDFScholar
2019

VERI-Wild: A Large Dataset and a New Method for Vehicle Re-Identification in the Wild

CVPR 2019poster

Vehicle Re-identification (ReID) is of great significance to the intelligent transportation and public security. However, many challenging issues of Vehicle ReID in real-world scenarios have not been fully investigated, e.g., the high viewpoint variations, extreme illumination conditions, complex ba…

Cited by 352PDFScholar