← Search

Jiaxuan Sun

2 accepted papers

2026

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition

CVPR 2026

Open-domain visual entity recognition (VER) seeks to associate images with entities in encyclopedic knowledge bases such as Wikipedia. Recent generative methods tailored for VER demonstrate strong performance but incur high computational costs, limiting their scalability and practical deployment. In

Cited by 0SourcecodeScholar
2025

NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation

NeurIPS 2025poster

Reinforcement learning (RL) has shown promise in enhancing the general Chain-of-Thought (CoT) reasoning capabilities of multimodal large language models (MLLMs). However, when applied to improve general CoT reasoning, existing RL frameworks often struggle to generalize beyond the training distributi…

Cited by 0SourceScholar