← Search

Jiashen Hua

3 accepted papers

2026

Illuminating Visual Identity in Universal Multimodal Embeddings

CVPR 2026

Universal Multimodal Embeddings (UMEs) aim to unify various modalities and tasks into a shared representation space. In recent years, this field has witnessed substantial progress driven by the development of Multimodal Large Language Models (MLLMs). However, a crucial capability, visual identity di

Cited by 0SourcecodeScholar
2026

Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs

ICLR 2026oral

Reasoning has emerged as a key capability of large language models. In linguistic tasks, this capability can be enhanced by self-improving techniques that refine reasoning paths for subsequent fine-tuning. However, extending these language-based self-improving approaches to vision language models (V…

Cited by 0SourcecodeScholar
2022

Online Convolutional Re-Parameterization

CVPR 2022poster

Structural re-parameterization has drawn increasing attention in various computer vision tasks. It aims at improving the performance of deep models without introducing any inference-time cost. Though efficient during inference, such models rely heavily on the complicated training-time blocks to achi…

Cited by 91PDFcodeScholar