← Search

Zixuan Song

2 accepted papers

2026

GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization

CVPR 2026

Cross-view geo-localization infers a location by retrieving geo-tagged reference images that visually correspond to a query image. However, the traditional satellite-centric paradigm limits robustness when high-resolution or up-to-date satellite imagery is unavailable. It further underexploits compl

Cited by 0SourcecodeScholar
2025

Towards General Continuous Memory for Vision-Language Models

NeurIPS 2025poster

Language models (LMs) and their extension, vision-language models (VLMs), have achieved remarkable performance across various tasks. However, they still struggle with complex reasoning tasks that require multimodal or multilingual real world knowledge. To support such capabilities, an external memor…

Cited by 0SourcecodeScholar