← Search

Na Zheng

5 accepted papers

2026

Breaking the Regional Perception Bottleneck of Multimodal Large Language Models via External Reasoning Framework

CVPR 2026

High-quality pixel-level responses remain a major bottleneck for multimodal large language models (MLLMs) in regional perception. Existing approaches generally attach regression decoders to MLLM features, achieving strong grounding performance but compromising end-to-end design and increasing traini

Cited by 0SourceScholar
2026

TTOM: Test-Time Optimization and Memorization for Compositional Video Generation

ICLR 2026poster

Video Foundation Models (VFMs) exhibit remarkable visual generation performance, but struggle in compositional scenarios (\eg, motion, numeracy, and spatial relation). In this work, we introduce **Test-Time Optimization and Memorization (TTOM)**, a training-free framework that aligns VFM outputs wi…

Cited by 0SourceScholar
2025

Language-Assisted Debiasing and Smoothing for Foundation Model-Based Semi-Supervised Learning

CVPR 2025poster

Recent studies have focused on introducing pre-trained foundation models into semi-supervised learning (SSL) tasks. Nevertheless, these foundation models can exhibit biases toward different classes and tend to generate imbalanced pseudo-labels for SSL. Thus, efforts have been made to introduce the l…

2025

Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization

NeurIPS 2025poster

Direct Preference Optimization (DPO) has emerged as an effective approach for mitigating hallucination in Multimodal Large Language Models (MLLMs). Although existing methods have achieved significant progress by utilizing vision-oriented contrastive objectives for enhancing MLLMs' attention to visua…

Cited by 0SourceScholar
2024

VK-G2T: Vision and Context Knowledge Enhanced Gloss2text

ICASSP 2024accepted

Existing sign language translation methods follow a two-stage pipeline: first converting the sign language video to a gloss sequence (i.e., Sign2Gloss) and then translating the generated gloss sequence into a spoken language sentence (i.e., Gloss2Text). While previous studies have focused on boostin…

Cited by 0SourceScholar