← Search

Borui Jiang

6 accepted papers

2026

DREAM: Document Recognition with Explicit Adaptive Memory

CVPR 2026

Large multimodal models (LMMs) have shown promising performance for various document recognition tasks. However, LMMs adopt implicit modeling, and the parameters lack interpretability. Inspired by recent advances in human memory and learning research, we propose an explicit multiscale prototype memo

Cited by 0SourcecodeScholar
2026

PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models

ICLR 2026poster

Multimodal large language models (MLLMs) have achieved strong performance on vision-language tasks, yet often suffer from inefficiencies due to redundant visual tokens. Existing token merging methods reduce sequence length but frequently disrupt spatial layouts and temporal continuity by disregardin…

Cited by 0SourcecodeScholar
2025

Single Domain Generalization for Few-Shot Counting via Universal Representation Matching

CVPR 2025poster

Few-shot counting estimates the number of target objects in an image using only a few annotated exemplars. However, domain shift severely hinders existing methods to generalize to unseen scenarios. This falls into the realm of single domain generalization that remains unexplored in few-shot countin…

2018

Acquisition of Localization Confidence for Accurate Object Detection

ECCV 2018poster

Modern CNN-based object detectors rely on bounding box regression and non-maximum suppression to localize objects. While the probabilities for class labels naturally reflect classification confidence, localization confidence is absent. This makes properly localized bounding boxes degenerate during i…