← Search

Bingyi Cao

10 accepted papers

2026

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

CVPR 2026

Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classification, retrieval, segmentation and depth prediction. However, a fundamental capability that these models still struggle with is aligning dense patch r

Cited by 0SourcecodeScholar
2025

Learning Visual Composition through Improved Semantic Guidance

CVPR 2025poster

Visual imagery does not consist of solitary objects, but in-stead reflects the composition of a multitude of fluid con-cepts. While there have been great advances in visual repre-sentation learning, such advances have focused on buildingbetter representations for a small number of discrete objectsbe…

Cited by 0SourcePDFScholar
2025

TIPS: Text-Image Pretraining with Spatial awareness

ICLR 2025poster

While image-text representation learning has become very popular in recent years, existing models tend to lack spatial awareness and have limited direct applicability for dense understanding tasks. For this reason, self-supervised image-only pretraining is still the go-to method for many dense visio…

2024

OmniGlue: Generalizable Feature Matching with Foundation Model Guidance

CVPR 2024poster

The image matching field has been witnessing a continuous emergence of novel learnable feature matching techniques with ever-improving performance on conventional benchmarks. However our investigation shows that despite these gains their potential for real-world applications is restricted by their l…

2023

Cooperative LiDAR Localization and Mapping for V2X Connected Autonomous Vehicles

IROS 2023poster

Cooperative Simultaneous Localization and Mapping (C-SLAM) is an active research topic in mobile robotics. However, its application in the field of autonomous driving is rare. While the advent of Vehicle-to-Everything (V2X) communication has empowered Connected Autonomous Vehicles (CAV) to exchange…

Cited by 7SourceScholar
2023

Global Features are All You Need for Image Retrieval and Reranking

ICCV 2023poster

Image retrieval systems conventionally use a two-stage paradigm, leveraging global features for initial retrieval and local features for reranking. However, the scalability of this method is often limited due to the significant storage and computation cost incurred by local feature matching in the r…

Cited by 47PDFcodeScholar
2023

Towards Universal Image Embeddings: A Large-Scale Dataset and Challenge for Generic Image Representations

ICCV 2023poster

Fine-grained and instance-level recognition methods are commonly trained and evaluated on specific domains, in a model per domain scenario. Such an approach, however, is impractical in real large-scale applications. In this work, we address the problem of universal image embedding, where a single un…

Cited by 17PDFScholar
2021

LiDAR-Based Object-Level SLAM for Autonomous Vehicles

IROS 2021poster

Simultaneous localization and mapping (SLAM) is an essential technique for autonomous driving. Recently, combining image recognition technology to generate semantically meaningful maps has become a new trend in visual SLAM research. However, in the field of LiDAR SLAM, this potential has not been fu…

Cited by 15SourceScholar
2020

Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and Retrieval

CVPR 2020oral

While image retrieval and instance recognition techniques are progressing rapidly, there is a need for challenging datasets to accurately measure their performance -- while posing novel challenges that are relevant for practical applications. We introduce the Google Landmarks Dataset v2 (GLDv2), a n…

Cited by 438PDFcodeScholar