← Search

Xueying Jiang

5 accepted papers

2024

LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained Descriptors

ICLR 2024poster

Inspired by the outstanding zero-shot capability of vision language models (VLMs) in image classification tasks, open-vocabulary object detection has attracted increasing interest by distilling the broad VLM knowledge into detector training. However, most existing open-vocabulary detectors learn by…

Cited by 26SourcePDFScholar
2024

MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked Autoencoders

NeurIPS 2024poster

Monocular 3D object detection aims for precise 3D localization and identification of objects from a single-view image. Despite its recent progress, it often struggles while handling pervasive object occlusions that tend to complicate and degrade the prediction of object dimensions, depths, and orien…

Cited by 4SourcePDFScholar
2024

Weakly Supervised Monocular 3D Detection with a Single-View Image

CVPR 2024poster

Monocular 3D detection (M3D) aims for precise 3D object localization from a single-view image which usually involves labor-intensive annotation of 3D detection boxes. Weakly supervised M3D has recently been studied to obviate the 3D annotation process by leveraging many existing 2D annotations but i…

Cited by 7SourcePDFScholar
2023

Black-Box Unsupervised Domain Adaptation with Bi-Directional Atkinson-Shiffrin Memory

ICCV 2023poster

Black-box unsupervised domain adaptation (UDA) learns with source predictions of target data without accessing either source data or source models during training, and it has clear superiority in data privacy and flexibility in target network selection. However, the source predictions of target data…

Cited by 19PDFcodeScholar
2023

Domain Generalization via Balancing Training Difficulty and Model Capability

ICCV 2023poster

Domain generalization (DG) aims to learn domaingeneralizable models from one or multiple source domains that can perform well in unseen target domains. Despite its recent progress, most existing work suffers from the misalignment between the difficulty level of training samples and the capability of…

Cited by 18PDFScholar