← Search

Wenzheng Song

4 accepted papers

2026

Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMs

AAAI 2026technical

Recent breakthroughs in reasoning models have markedly advanced the reasoning capabilities of large language models, particularly via training on tasks with verifiable rewards. Yet, a significant gap persists in their adaptation to real-world multimodal scenarios, most notably, vision-language tasks

Cited by 0SourcePDFScholar
2026

When Diffusion Language Models Hesitate: Detecting and Correcting Visual Hallucinations via Confidence Fluctuation

ICML 2026poster

Multi-modal Diffusion Language Models (MDLMs) have emerged as a powerful alternative to autoregressive models in vision-language understanding, offering advantages in bidirectional context modeling and parallel decoding. However, existing MDLMs suffer from severe visual hallucinations due to the sta…

Cited by 0SourceScholar
2024

Globalizing Local Features: Image Retrieval Using Shared Local Features with Pose Estimation for Faster Visual Localization

ICRA 2024poster

Visual localization is an important sub-task in SfM and visual SLAM that involves estimating a 6-DoF camera pose for an input query image relative to a given 3D model of the environment. The most accurate approach is a hierarchical one that splits the task into two stages: image retrieval and camera…

Cited by 0SourceScholar
2021

Matching in the Dark: A Dataset for Matching Image Pairs of Low-Light Scenes

ICCV 2021poster

This paper considers matching images of low-light scenes, aiming to widen the frontier of SfM and visual SLAM applications. Recent image sensors can record the brightness of scenes with more than eight-bit precision, available in their RAW-format image. We are interested in making full use of such h…

Cited by 21PDFcodeScholar