← Search

Jiarui Cai

6 accepted papers

2026

Decoupling Vision and Language: Codebook Anchored Visual Adaptation

CVPR 2026

Large Vision-Language Models (LVLMs) use their vision encoders to translate images into representations for downstream reasoning, but the encoders often underperform in domain-specific visual tasks such as medical image diagnosis or fine-grained classification, where representation errors can cascad

Cited by 0SourceScholar
2026

Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes

CVPR 2026

We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text-instructed spatial transformation of objects within scenes. Spatially manipulating objects in a scene through natural language poses a challenge for multimodal generation systems. While existing text-based manip

Cited by 0SourcecodeScholar
2024

Hyperbolic Learning with Synthetic Captions for Open-World Detection

CVPR 2024poster

Open-world detection poses significant challenges as it requires the detection of any object using either object class labels or free-form texts. Existing related works often use large-scale manual annotated caption datasets for training which are extremely expensive to collect. Instead we propose t…

Cited by 6SourcePDFScholar
2022

LUNA: Localizing Unfamiliarity Near Acquaintance for Open-Set Long-Tailed Recognition

AAAI 2022technical

The predefined artificially-balanced training classes in object recognition have limited capability in modeling real-world scenarios where objects are imbalanced-distributed with unknown classes. In this paper, we discuss a promising solution to the Open-set Long-Tailed Recognition (OLTR) task utili…

Cited by 14SourcePDFScholar
2021

ACE: Ally Complementary Experts for Solving Long-Tailed Recognition in One-Shot

ICCV 2021poster

One-stage long-tailed recognition methods improve the overall performance in a "seesaw" manner, i.e., either sacrifice the head's accuracy for better tail classification or elevate the head's accuracy even higher but ignore the tail. Existing algorithms bypass such trade-off by a multi-stage trainin…

Cited by 185PDFcodeScholar