← Search

Zhenhua Chai

9 accepted papers

2024

Investigating Compositional Challenges in Vision-Language Models for Visual Grounding

CVPR 2024highlight

Pre-trained vision-language models (VLMs) have achieved high performance on various downstream tasks which have been widely used for visual grounding tasks in a weakly supervised manner. However despite the performance gains contributed by large vision and language pre-training we find that state-of…

2024

MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation

ACL 2024long

Embodied agents equipped with GPT as their brain have exhibited extraordinary decision-making and generalization abilities across various tasks. However, existing zero-shot agents for vision-and-language navigation (VLN) only prompt the GPT-4 to select potential locations within localized environmen…

Cited by 30SourcePDFScholar
2023

DRKF: Distilled Rotated Kernel Fusion for Efficient Rotation Invariant Descriptors in Local Feature Matching

IROS 2023poster

The performance of local feature descriptors degrades in the presence of large rotation variations. To address this issue, we present an efficient approach to learning rotation invariant descriptors. Specifically, we propose Rotated Kernel Fusion (RKF) which imposes rotations on the convolution kern…

Cited by 3SourceScholar
2023

Divide and Adapt: Active Domain Adaptation via Customized Learning

CVPR 2023highlight

Active domain adaptation (ADA) aims to improve the model adaptation performance by incorporating the active learning (AL) techniques to label a maximally-informative subset of target samples. Conventional AL methods do not consider the existence of domain shift, and hence, fail to identify the truly…

2022

Compressing Models With Few Samples: Mimicking Then Replacing

CVPR 2022poster

Few-sample compression aims to compress a big redundant model into a small compact one with only few samples. If we fine-tune models with these limited few samples directly, models will be vulnerable to overfit and learn almost nothing. Hence, previous methods optimize the compressed model layer-by-…

Cited by 13PDFcodeScholar
2021

Feature Decomposition and Reconstruction Learning for Effective Facial Expression Recognition

CVPR 2021poster

In this paper, we propose a novel Feature Decomposition and Reconstruction Learning (FDRL) method for effective facial expression recognition. We view the expression information as the combination of the shared information (expression similarities) across different expressions and the unique informa…

Cited by 218PDFScholar
2021

Rethinking BiSeNet for Real-Time Semantic Segmentation

CVPR 2021poster

BiSeNet has been proved to be a popular two-stream network for real-time segmentation. However, its principle of adding an extra path to encode spatial information is time-consuming, and the backbones borrowed from pretrained tasks, e.g., image classification, may be inefficient for image segmentati…

Cited by 814PDFcodeScholar
2021

Trash To Treasure: Harvesting OOD Data With Cross-Modal Matching for Open-Set Semi-Supervised Learning

ICCV 2021poster

Open-set semi-supervised learning (open-set SSL) investigates a challenging but practical scenario where out-of-distribution (OOD) samples are contained in the unlabeled data. While the mainstream technique seeks to completely filter out the OOD samples for semi-supervised learning (SSL), we propose…

Cited by 76PDFScholar