← Search

Rui Cai

8 accepted papers

2025

Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval

AAAI 2025technical

Existing cross-modal retrieval methods typically rely on large-scale vision-language pair data. This makes it challenging to efficiently develop a cross-modal retrieval model for under-resourced languages of interest. Therefore, Cross-lingual Cross-modal Retrieval (CCR), which aims to align vision a…

2021

AVP-Loc: Surround View Localization and Relocalization Based on HD Vector Map for Automated Valet Parking

IROS 2021poster

Localization is a crucial prerequisite for automated valet parking, in which a vehicle is required to navigate itself in a GPS-denied parking lot. Traditional visual localization methods usually build a feature map and use it for future localizations. However, the feature map is not robust to change…

Cited by 13SourceScholar
2021

DT-Loc: Monocular Visual Localization on HD Vector Map Using Distance Transforms of 2D Semantic Detections

IROS 2021poster

Localizing a vehicle on a prebuilt HD vector map is a prerequisite for many autonomous driving applications. Existing visual localization approaches usually require a separate local feature layer to function. The separate localization layer suffers from the robustness issue inherited from the local…

Cited by 11SourceScholar
2021

Robust LiDAR Localization on an HD Vector Map without a Separate Localization Layer

IROS 2021poster

Many autonomous driving applications nowadays come along with a prebuilt vector map for routing and planning purposes. In order to localize on this map, traditional LiDAR localization methods usually require a separate localization layer to function. On one hand, the separate layer occupies large st…

Cited by 11SourceScholar
2017

Locality-Sensitive Deconvolution Networks With Gated Fusion for RGB-D Indoor Semantic Segmentation

CVPR 2017poster

This paper focuses on indoor semantic segmentation using RGB-D data. Although the commonly used deconvolution networks (DeconvNet) have achieved impressive results on this task, we find there is still room for improvements in two aspects. One is about the boundary segmentation. DeconvNet aggregates…

Cited by 273PDFScholar
2016

Joint Multiview Segmentation and Localization of RGB-D Images Using Depth-Induced Silhouette Consistency

CVPR 2016poster

In this paper, we propose an RGB-D camera localization approach which takes an effective geometry constraint, i.e. silhouette consistency, into consideration. Unlike existing approaches which usually assume the silhouettes are provided, we consider more practical scenarios and generate the silhouett…

Cited by 7PDFScholar
2015

MeshStereo: A Global Stereo Model With Mesh Alignment Regularization for View Interpolation

ICCV 2015oral

We present a novel global stereo model designed for view interpolation. Unlike existing stereo models which only output a disparity map, our model is able to output a 3D triangular mesh, which can be directly used for view interpolation. To this aim, we partition the input stereo images into 2D tria…

Cited by 206PDFScholar
2015

Query Adaptive Similarity Measure for RGB-D Object Recognition

ICCV 2015poster

This paper studies the problem of improving the top-1 accuracy of RGB-D object recognition. Despite of the impressive top-5 accuracies achieved by existing methods, their top-1 accuracies are not very satisfactory. The reasons are in two-fold: (1) existing similarity measures are sensitive to object…

Cited by 18PDFScholar