← Search

Bairui Wang

5 accepted papers

2025

VITRIX-UniViTAR: Unified Vision Transformer with Native Resolution

NeurIPS 2025poster

Conventional Vision Transformer streamlines visual modeling by employing a uniform input resolution, which underestimates the inherent variability of natural visual data and incurs a cost in spatial-contextual fidelity. While preliminary explorations have superficially investigated native resolution…

Cited by 0SourceScholar
2019

Controllable Video Captioning With POS Sequence Guidance Based on Gated Fusion Network

ICCV 2019poster

In this paper, we propose to guide the video caption generation with Part-of-Speech (POS) information, based on a gated fusion of multiple representations of input videos. We construct a novel gated fusion network, with one particularly designed cross-gating (CG) block, to effectively encode and fus…

Cited by 232PDFcodeScholar